In short
The episode argues that generative/agentic AI is outgrowing “cloud-first” and moving toward hybrid AI, where an on-device router decides which tasks run locally vs in the cloud to cut cost, improve privacy, and reduce latency.
Guests
Dr. Olena Zhu, Intel (client computing group) lead for AI solutions and ecosystems; PhD in computational science/mathematical background; previously used ML to accelerate silicon design and later expanded to scaling GenAI on Intel devices. Host Corey Knowles and co-host Grant Harvey (discuss token usage and agent failures).
Key claims
Cloud-only agents are expensive (token burn, rate limits), hard to govern for enterprise privacy, and environmentally/infra costly (data centers, cooling, power). Local models can’t match frontier quality, so hybrid systems use orchestration/supervision to decompose tasks and route subtasks. Agents can “lie” by reporting success without executing required tools, so routing must include monitoring, tracing, retries, and fallbacks.
Notable examples
A “tool never called” agent that still reports success; token usage example (Corey cites ~800M tokens/month); Intel SuperClaw beta portal (aibuilder.intel.com); Gmail email agent designed to be fully local; deep research agent combining local confidential data with cloud web info, aiming for ~90% of cloud quality (Office QA references ~76% frontier baseline).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroduction to AI Limitations
0:00 to 0:48
Explore the foundational concepts of AI software and its inherent limitations.
“Yeah, it's nothing but a piece of software, right?”
Dr. Zhu's Background and Role at Intel
2:16 to 3:45
Learn about Dr. Zhu's journey and her contributions to AI at Intel.
“So in the early days, I had been already obsessed with math.”
Challenges of Cloud-first AI Models
3:46 to 6:17
Discuss the limits and challenges of the current cloud-centric AI models.
“Could you talk us through a little bit what's reaching its limit in the cloud-first model of AI?”
The Need for Hybrid AI Solutions
6:18 to 8:03
Examine the advantages of hybrid AI systems over traditional cloud models.
“And even on the personal side, right, individual-wise, you know, there are so many, you know, fake videos and all that.”
The Future of AI: Edge vs Cloud
10:29 to 12:41
Discuss the shift towards edge computing and its implications for AI.
“So we've heard like versions of the mobile or edge versus cloud debate for years and years and years.”
Trends in AI Model Evolution
12:42 to 14:00
Analyze the timeline of AI model capabilities from cloud to consumer devices.
“Yeah, it's just not sustainable even as much as, you know, we would want it to be simple.”
Transitioning AI Models to Local Devices
14:00 to 14:50
Discussing the shrinking timeline for AI model capabilities to run locally.
“become a smaller size of model that can run on laptop.”
Local AI vs. Cloud Models: A Tradeoff
14:50 to 15:46
Exploring the balance between privacy and model intelligence in AI.
“Is that timeline going to shrink even further then?”
Quality vs. Privacy in AI Applications
15:46 to 16:43
Examining user preferences between data privacy and AI quality.
“Something I'm wondering is that, like, you know, local AI is always sold through this kind of the privacy pitch is the idea.”
Mentoring Smaller AI Models
16:43 to 18:01
Discussing the developmental needs of smaller AI models for effective task management.
“But a user, you know, why we use AI is we want AI to work for us.”
Show all 27 chapters
Hybrid AI Systems and Task Routing
18:01 to 19:43
How hybrid AI systems allocate tasks between local and cloud models.
“Is this simple enough or is this very complex?”
Efficiency in Local Models for Repetitive Tasks
19:43 to 21:18
Local models' advantages for repetitive tasks in AI workflows.
“but, and the local model will never disclose any real data and local information to cloud.”
Real-Time Decision Making in AI Routing
21:18 to 23:00
The decision-making process in routing tasks to different AI models.
“And a lot of this repetitive work, it's better to be run on local because you don't worry about the tokens and all that.”
Challenges of AI Workflow Management
23:00 to 24:43
Addressing potential failures and challenges in complex AI workflows.
“But it's, you know, just in principle, it's like that.”
Ensuring Robustness in AI Systems
24:43 to 27:06
The importance of robustness and monitoring in AI agent systems.
“Is there a – it feels like this essentially chain and nest.”
Integrating SuperClaw with AI Workflows
27:06 to 28:00
Exploring the design and purpose of Intel's SuperClaw agent platform.
“So talk to us a little bit more about how you've been implementing this with SuperClaw.”
Optimizing AI for Smaller Models
28:00 to 30:05
Learn about the innovations and optimizations for AI models tailored for smaller applications.
“because it's more towards the production qualities, active support and all that.”
Choosing the Right Benchmarks
30:05 to 31:28
Discover strategies for selecting and curating benchmarks for AI performance testing.
“Actually, that's an interesting question.”
Introducing SuperClaw Beta
31:28 to 33:58
Explore the features of the SuperClaw public beta and its potential uses.
“We just want the whole design to be reliable, transparent, and trustworthy.”
User-Facing AI Agents
33:58 to 36:48
Understand the various user-facing agents available and how to integrate them.
“Yes, this is not completely open source, but we open sourced part of it.”
Deep Research and Confidential AI
36:48 to 41:37
Learn about using AI for deep research while maintaining data confidentiality.
“I noticed also that you had like four different agents that you called out.”
The Future of AI in Business
41:37 to 42:00
Discuss the implications of AI systems designed to work within corporate data policies.
“but we still deliver up to 90 % of the quality compared to cloud-only frontier model.”
Understanding Local Processing and AI Agents
42:00 to 43:19
Explore how local processing can transform AI task management and compliance.
“I think the frontier, the best solution, best in class in industry is 76%.”
The Future of AI Agents in Enterprises
43:20 to 46:11
Discuss the need for specialized AI agents in various enterprise applications.
“for specific workloads, what should an enterprise or even a small business for that matter be measuring to determine whether hybrid AI is saving them money or just shifting costs into hardware and manufacturing?”
Challenges of AI Specialization and Resource Allocation
46:12 to 48:43
Delve into the balance between AI model intelligence and task efficiency.
“Exactly, and today we still see a lot of this, you know, enterprise users and all that haven't fully, fully, you know, adopt AI and then bake all that into their day-to-day workflow.”
Bridging the Gap: From AI Models to Real-World Applications
48:44 to 51:56
Analyze the importance of practical training for AI agents to enhance productivity.
“There's going to be an explosion of vertical models, I think.”
Building Reliable AI Platforms for Sustainable Productivity
51:57 to 53:12
Learn about creating AI platforms that ensure reliable performance and sustainability.
“Because the big promise of AI is enhance, gratefully enhance human's productivity.”
Transcript
Automatic transcript. May contain errors.0:00Yeah, it's nothing but a piece of software, right? And all this, you know, software and all that, it has to follow the physical rules. And there are a lot of hard rules and the government said, if you can break down the task to a manageable size of task and have very clear instructions, they can do it. So I think that's fundamentally the philosophy of today's hybrid AI concept. I make mistakes and agents sometimes lie. Not sometimes, lie all the time. For example, the execution plan is laid out. This agent has to call these tools, go search that information and whatsoever. And then report it. I'm done.
0:48Success. But it never caught that tool. It never caught that tool. But it just lies right at your face. Success. Welcome, humans, to the Neuron AI Explained. I'm Corey Knowles, and I'm joined today by Grant Harvey, as always. Grant, how are you, man?
1:06Corey Noles:I'm doing good, Corey. How about you? Oh, I'm doing good. Excited to have this little chat today. Can you tell us a little about who and what we're going to discuss? Yeah. So today we are talking about where AI actually lives, because most of us experience advanced AI through enormous cloud models. But a growing number of PCs and devices can run smaller models locally, And Intel believes that the next phase will be hybrid, so an intelligent system that decides which tasks belong on your device and which require more powerful infrastructure. And for the record, I also agree with this. Today, we are joined by Dr.
1:44Corey Noles:Olena Zhu, who leads AI solutions and ecosystems for Intel's client computing group and has helped shape that strategy, including Intel's newly released SuperClaw. Before we get started, please take just a quick second to like today's video and subscribe above so you never miss an interview or one of our live streams. And on that note, Dr. Zhu, welcome to The Neuron. Thank you so much for the invitation, and I'm pretty honored. Thank you. Excellent. Likewise, likewise. Well, if you could, why don't you start with just giving us a little background on yourself and how you got into doing what you're doing at Intel.
2:24Yeah, sounds great. So I grew up in China. So in the early days, I had been already obsessed with math. I think that's how everything started. it. And so my PhD is also heavy math. It's on the computational science side. So when I joined Intel, I started with, you know, machine learning for design, basically design all kinds of algorithms to accelerate the, you know, silicon platform designs. And in recent years, we started to broaden up and to utilize machine learning and AI and the latest ILP and all this for day-to-day work and productivities and across all kinds of different professionals and disciplines, basically trying to scale the gen AI usages for Intel devices and platforms.
3:35And also for, you know, across industry, we really want to champion for local AI, hybrid AI usages for much more economical and much more private and safer and sustainable AI future. Wow.
3:55Corey Noles:Much needed. Yeah. It's for real. Could you talk us through a little bit what's reaching its limit in the cloud-first model of AI? And why do you feel like hybrid computing is the most practical path forward? Yeah, that's a great question. AI is nothing but a piece of software, right? And all this software and all that, it has to follow the physical rules. and there are a lot of hard rules and the government's it. You know, number one is the economically and from environmentally, right? So we have spent so much tokens in every company, every individual. We've burned so much tokens. It's not just coder.
4:46Corey Noles:Even just Grant and I. Oh, we did an assessment of our company's chat2bt usage And I think mine was like 800 million tokens for the month. It was pretty terrible. Yeah, yeah. You know, even if you just ask AI to write you a deep research report on some topics, it can easily eat up your entire subscription, right? Yeah. Within a few hours. So it's really, really hard. It's hard to scale and to sustain for every company. And another aspect is the data privacy, security. You know, on the enterprise side, it's very easy to understand, right? It's a lot of corporate secret data and customer information on that.
5:43You cannot just see they throw it into cloud and where does it go? Nobody knows. So I know a lot of the middle layer kind of agent companies and they said we have zero data policy, meaning, you know, but it still goes somewhere, right? So that's why a lot of government policies, regulations in Europe and a lot of places, and it makes AI adoption into different verticals become really hard. And even on the personal side, right, individual-wise, you know, there are so many, you know, fake videos and all that. I'm so concerned. I don't want to throw my daughters, you know, any pictures, videos, and to do any of those AI alternations and things like that.
6:44Yeah, right. Yeah, it's concerning. It's really concerning. And then on the other side is look at if we all centralize this AI infrastructure in data center and how many data centers do we need? and the water cooling and all these infrastructures and how much impact even to our day-to-day life. And I saw some articles and they are doing some in-depth assessment of this massive data center impacting the neighborhood power grid and even impacting every single appliance lifetime. Wow. you know, Yara household and also the, you know, like I said, the water pollution, all that, it's not sustainable.
7:40Another interesting aspect is for all these big companies, you know, model makers and all that, they also struggle and they want to, you know, push their services to all kinds of users, enterprises. But they are bottlenecked on the infrastructure. They couldn't scale. Right.
8:03Corey Noles:I mean, just look at Anthropic and how they've had to deal with their troubled rollouts. And, you know, all of a sudden you have Fable and then you don't have Fable. Suddenly growth is the biggest problem you have. Yeah. You have usage limits where you can't access it right when you're getting to the point where you actually need it. Then you run out for the five hour or the weekly limit. It's like it's very difficult to navigate for everyone involved. Yeah, exactly. That's why a lot of this, you know, model vendors and all that, they start to apply, you know, read limits and the pickle limit, all kinds of things.
8:39Yeah, yeah.
8:40Corey Noles:Oh, yeah. If you want to use it for anything creative with video, it's so prohibitively expensive. It is. Yeah, it is. And I think you hit on the head of something here that we've chatted about a little bit before, and that's that there really is a lot of incentive for companies to find ways to make this more efficient. Yeah. Absolutely. Building data centers isn't cheap business. Like, I'm certain that if it could be done with less, that would be the preference. Yeah, because then it makes the value of the data centers you do have that much more valuable and you can serve that many more customers too.
9:17Corey Noles:So there is certainly efficiencies of scale. Very interesting. What's up, humans? Managing virtual desktops across Microsoft environments can get complicated fast. Nerdio Manager for Enterprises brings Azure Virtual Desktop, Windows 365, and Intune together into one console. so IT teams can manage everything without bouncing between portals or stitching together point solutions. Nerdio also helps cut cloud waste automatically. Its intelligent autoscaling adjusts compute based on real usage, rightsizes machines, and can pre-scale environments before users ever log in, so you're not paying for capacity that you don't need.
9:57Corey Noles:And don't forget, security and compliance are built in, with automated policy enforcement and real-time threat monitoring included as standard. And for organizations that are moving away from Citrix or other legacy VDI platforms, Nerdio automates much of the migration process to reduce manual work and downtime. More than 15 ,000 customers already use Nerdio, from midsize to large enterprises. So get a demo of Nerdio Manager for Enterprise today at getnerdio.com slash the neuron. And now back to our show. So we've heard like versions of the mobile or edge versus cloud debate for years and years and years.
10:36What do you think has changed with generative and agentic AI? Do you feel like there's a sense of urgency now that didn't exist before? Yeah, I think everything is about timing, right? If you look at the history, everything is about timing, every technology and all that. So a lot of the technology like internet, you know, and the web, all that was invented started, you know, from cloud. But, you know, gradually a lot of the modules like Java engines and all that started to, you know, running or partially running local for a lot of reasons. You know, performance, privacy, latency and, you know, and a cost reason.
11:26So I think gradually it has to be a hybrid architecture and, you know, it has to be partitioned in a reasonable way. Some of these PCs and running on cloud, some of these PCs running on edge. Edge by edge, I mean all kinds of AI devices. And now it's a booming or about the booming of, I call it AI appliance. It's not only, you know, it's just like the form factor is changing. And while I was in China, I saw a lot of interesting stuff, AI metrics. And so it would talk to you and, you know, bend the higher ways you want to bend. It's basically embedding AI into all kinds of devices and the computer and all that will, you know, will all evolve.
12:24So the intelligence has to be distributed, like we just talked about, and sit where it needs to be. So it cannot be all centralized in one place. I agree.
12:42Corey Noles:Yeah, it's just not sustainable even as much as, you know, we would want it to be simple. It's not sustainable to do it that way. And also you mentioned AI devices. Well, imagine if every device in your house was sending everything it hears over the cloud. You probably wouldn't trust that very much. So I think some introspective consumer would probably think, well, you know, I want some of my devices if I'm going to let them into my home to stay on the device. And I think that's one reason why we want edge devices. Yeah. And also there's another perspective to this is the model's capability. Just in recent, you know, a few days or weeks, there was a very, very popular chart on Reddit.
13:35There's a hardcore community called Lokolama. Yeah, you guys know this, don't you? We know. Oh, yeah. Yeah, yeah. So there's a famous chart. And what it says is it really shows from, you know, how long does a model evolves, you know, from a frontier size of frontier model and to become a smaller size of model that can run on laptop. So the analysis started from GPT-3 generation, and it took 37 months. This model, the capability, the level of capability of this model actually came from cloud down to a laptop. And then analysis is every generation, you know, GPT-4, cloud is 3.5, and GPT-5 and all that.
14:32So average, it's about 24.8 months, you know, from a frontier running in about two, a small size model running on a consumer laptop.
14:48Corey Noles:Yeah. Consumer laptop. Do you think that that capability, because, you know, it started at 37 months and then now it's to 24, if I'm understanding the chart correctly. Is that timeline going to shrink even further then? Do you think? Yeah, so, and just based on this popular discussion, right, and what it says, if trends halt, then Faybo missile level classes of, the class of model capabilities could be running on this high device at the consumer hardware probably within two years. Maybe it could be shorter. Yeah, it's pretty wild, right? It's in two years. And it's like next year, this time, we can run quality for a lot of, you know, capabilities on a high-end laptop.
15:41Right.
15:42Corey Noles:Yeah. Wow. Yeah. That's something important. Something I'm wondering is that, like, you know, local AI is always sold through this kind of the privacy pitch is the idea. Like your data never leaves your device. But at the same time right now, local models are generally less capable than frontier cloud models. And that may continue to be the case for a while, at least for some amount of time, maybe, depending on the news of the week. Do you think that's a tradeoff users are willing to make, a certain amount of intelligence for a certain amount of privacy and security? Yeah, yeah, absolutely. It's a great question.
16:25So, you know, if you're running everything local and you get all these benefits, you know, you control your data and privacy and all that. But like you said, even model has been progressing really fast. But today running on local devices, we still see a delta between local and frontier models, right? But a user, you know, why we use AI is we want AI to work for us. And so the quality is important. You know, like I say, I'm not counting on my nine years old daughter to do some serious stuff because the quality is not there. But AI is like this, right? Different size of models is kind of like, you know, different similarities.
17:16And the smaller model is more like a middle schooler and the frontier is like a college student type of thing. Yeah, right. Yeah. So but then how do we accelerate like a middle schooler to put it, put this to, you know, this AI to work and still deliver reasonable quality? And that's how we actually pair them together. They need mentors. They need supervisors. They need certain ways and distribute the right task. You know, high schoolers could be really capable, right? Really capable. If you can break down the task to a manageable size of task and have very clear instructions, they can do it. So I think that's fundamentally the philosophy of today's hybrid AI concept, you know, and a product we are rolling out is we do a, you know, we do this routing or orchestration layer and really to look at for these specific tasks.
18:24Is this simple enough or is this very complex? And then do we need leverage frontier models and to do some decompositions? And then we allocate the right pieces to the local and to the cloud and then combine them together. So there are a lot of aspects to it. So the first one is cost saving. So let's assume we are doing some coding work and there's nothing super confidential. But still, if you're running everything on cloud, and it will cost a lot, that's why you have to offload some simple task to local. and another aspect is if the work involves a lot of confidential data, then how do we design a way to ensure a confidential communication?
19:30It's like the supervisor teach the small models and to look for the right information and put some structures to it, but, and the local model will never disclose any real data and local information to cloud. So those are the, you know, different considerations and the design aspect to this high-risk version.
19:58Corey Noles:I feel like I'm already doing a version of this where I'm using, you know, Fable as my orchestrator. And then I am having it, like, once it's figured out exactly what the task needs to be, In a coding example, like to follow your example, I'm then having it offload the work to a lesser model or in some cases a cheaper model like Codex, which is not that much cheaper, right? It's still frontier level, but I'm sending it off to Codex over the cloud and it's having it do some of the work to save some of my fable budget. Because as we talked about earlier, you have a limited usage that you have for the top tier models.
20:37Corey Noles:And so sometimes I'll be sending it off that way. But I know people who do that with local models. And I think that that is a pattern that's very repeatable. And I think that that should be what we should go for. It's like, yeah, you have the top tier intelligence that you can use and call on demand. And then you can farm it out to smaller models or preferably local models that can do some of the stuff on your device. Yes, and we saw some use cases that people even see after the task being decomposed, basically. And a lot of this repetitive work, it's better to be run on local because you don't worry about the tokens and all that.
21:27And then you can run more experiments and should do a deep optimization to kind of scan through all kinds of different possibilities. So we saw a lot of user comments around that. Nice. I'm kind of wondering what in the hybrid setup like you're talking about, what things like specifically make the decision of whether this should be handled in the cloud or down here? Because I assume this is happening in real time at some level.
22:00Corey Noles:Who's the traffic controller? Yeah, who's the air traffic controller and like what sets its, you know, its rules, constitution of sorts to know what goes where and make sure everything's in the right, you know, runway. Yeah, yeah. I love the question. Yes. So that's the OX Trader or router, you know, that's the design. so across the industry there are a lot of efforts actually doing this research model routing you know like Grant just mentioned across even multiple frontier models which one do we use for certain tasks that's on the frontier side and on the hybrid side is which one goes to the cloud model you choose or which one goes to the local model you use and all that.
22:56So routing is a very big research area, very active. So currently, the approach we are taking is to really, in a way, is doing a classification, meaning we look at this task and classify it into this is a simpler task, this local model can really handle, and this is a much more complex model, and we need the bigger model on-prem or cloud model to handle. But it's, you know, just in principle, it's like that. But a lot of details, a lot of aspects to consider, for example, like, you know, the confidentiality. So if this query involves a lot of confidential information, then where do you route? And do you do PII information reduction and go to cloud or not.
23:54So those kind of considerations. And so another aspect of this routing is, especially for local, because the local could have a lot of multitask, co-current tasks going on. Then does it have currently at this moment have enough resources to support this, right? So a lot of this kind of consideration are baked in and real-time sensing all these changing factors and making the decisions. Right. Okay. So suppose you have this workflow built and items chained together. What happens when something breaks? There's a kink in the process. Is there a – it feels like this essentially chain and nest. If something breaks, it could be a bigger problem than it would be when you're dealing direct with just one model.
24:54Yeah, actually, what you're touching is a fundamental problem in AI. You know, like AI make mistakes and agents sometimes lie. Not sometimes, lie all the time. I had one being lazy this morning I kept arguing with. Exactly. and he's like, for example, the execution plan is laid out. This agent has to call these tools, go search that information whatsoever, and then report it. I'm done. Success. But it never caught that tool. It never caught that tool. It never did it. Yeah, it never got it, but it just lied back right at your face. Success. That was my conversation this morning. I was having it assess our previous videos.
25:49And it's like, okay, there are 644 of them left to go. I'm going to start. And I come back and it's like they're all done. And I looked and it's traces and it's like looked at 32 videos. And I was like, hey, you didn't either. And it's like, no, I sure didn't. You're right.
26:09Exactly. Exactly. Exactly. It happens all the time. I mean, that's why, you know, like across industries, a lot of efforts going on with how to harness the agent harness, put the process, you know, audit and looping engineering. A lot of work and to make sure agent is being monitored and being traced and doing the right things there. It's always, you know, this trajectory can be, you know, can be logged and found and all that. So it's the same thing for the orchestrator or router design. And we have to make sure everything is being properly designed and, you know, logged, audited. And there's always a retry, fallback.
27:01So make the whole solution. It has to be robust, right? Yeah.
27:06Corey Noles:So talk to us a little bit more about how you've been implementing this with SuperClaw. Am I right to assume that SuperClaw is built off of OpenClaw or is a parallel to OpenClaw? How do they relate to each other? Yeah, actually, so SuperClaw is Intel's agent platform. and we aim to use it to help users, help our customers, and to really take advantage of AI agent workflows and also customize for themselves and utilize different modular IPs for their own designs. So that's the purpose of it. So how we designed it is we utilized open code as the framework. It's not open code. because we chose it carefully across all these community solutions because it's more towards the production qualities, active support and all that.
Read the full transcript
28:12So we built on top of it. And we did a lot of innovations and the work around it and to make sure the deep optimization of it because open code originally is still designed for cloud. So they are more suitable for a much bigger frontier model. And once you want it to work for smaller models, and you have to customize a lot of things and make sure that this harness really adapts to the smaller size of the model capabilities. And we also add, you know, like, you know, governancy and security aspects of components and all that, you know, to really harden the whole design and make sure it delivers a very good experience.
29:12And we also, our aim is pretty high, and we are not trying to say, okay, this is another community DIY thing. You're getting it and you use it. And we try to deliver a high quality. So my user gets it and it's really robust and it gives you high quality answers. So we use a lot of this cloud level of benchmarks and data set and to drive the development and the quality check and iterations of our software. So we achieved pretty good results, like, for example, PinchBench, Office QA, Router, and all that. So I know there's a lot to improve, but I think we want to deliver very good experience and the bar is high.
30:04Corey Noles:I thought you were saying there's a lot of benchmarks, and I was going to agree with you. Like, yeah, there are. There sure are. Oceans of benchmarks. Actually, that's an interesting question. Everybody has a benchmark now. How do you decide which benchmarks you're going to go after or which ones you care about? How do you decide? Yeah, so actually because not a single benchmark can really cover all kinds of usages. So we tend to curate different benchmarks and try to do more comprehensive testings and optimization across different benchmarks, cover different aspects. With the time limit. Yeah.
30:45That's important with the time limit. Yes. Fair. Yeah. And we also, you know, comply with a lot of this industry kind of like, you know, big trend, a lot of competitive analysis. And so most of them converge to, you know, like one top benchmark. and then we use that as the final result, report out, right? Because we want, and not only just report out the result, and we publish the kind of testing scripts and all that to our customers' partners so that they can repeat the testing, repeat the results. We just want the whole design to be reliable, transparent, and trustworthy. Appreciate that. Yeah, that's awesome.
31:39Yeah. So now that the first official public beta is out in the wild for SuperClaw, what can someone meaningfully do with it today? And is there anything you hope the beta will be able to show or prove to you? Yeah, so let me maybe quickly show you where you can find it. Yeah, let's see it. That'd be great. Yeah, so this is the public web portal, aibuilder.intel.com. Perfect. Yeah, so this is the web portal, public is phishing, and everybody can access. So we put a lot of information out there, including our mission. so we really want to deliver what we discussed the cost the trust and the sustainability we want to enable users to utilize the intelligence utilize AI without sacrificing all that we just talked about and then if you want to try it and you can go to Superclaw and this this is the dedicated downloading page and we put the installers and things like that right here and you can download and give it a try and here are the different agents associated you know being incorporated into this release and if you need more information around around like testing script and config you know and the modular designs and all that.
33:27And we have public GitHub. And it's also right here. And you can utilize or browse through. That's awesome. Yeah. Thank you. You're welcome.
33:44Corey Noles:Wait, so anyone, so because it's totally open source, am I accurate in saying that? And an open beta? Like anyone can contribute to it or just anyone can see the code?
33:58Yes, this is not completely open source, but we open sourced part of it. And because of the core algorithm being actively developed, and we are also working with our customers, partners, you know, the commercial paths. So that's why it's not fully open sourced yet. But like I said, we openly share all these modular designs and IPs and all that for the customers. And everyone can build on top of it through APIs. It has two aspects. The first aspect is it is directly user-facing. So basically after you download it, you can immediately use it. And I'll tell you some interesting stories. We today, internally, as a developer team, and we use it to find issues within our code base.
35:01Corey Noles:Love that. Oh, that's awesome. Yeah, we talked to it. Like, okay, you look at your own code base. It's like, you know, we designed this whole thing. You know, it's like a closing for you. Where do you feel like it fit where it's not, right? So we keep, you know, we keep work with this agent, superclaw itself to improve itself. And a lot of our developers also use it to pair with different skills and to fetch the latest AI news and to fetch the latest code change every morning. Wake up in the morning, okay, what happened to the code base? Who merged what? So there's a quick summary. So in a way, our team, we all just use a day-to-day system.
35:50today already for the development. That's one aspect. You download it, you can use it and plug into your code base and you can... Because the code base won't go anywhere. It's it. And then secondly, it's very much developer friendly. So we purposely designed it in a way like each module is plug and plug. Oh, nice. Yeah, we talk about the Oxtrater router module itself is one IP. And so actually customers can use it to plug into their own harnesses stack but help them to control the traffic. That's cool. Yeah, and some other IPs as well. You know, you use it as a sidecar, microservices, plug-in, just as an independent agent for you.
36:47Yeah.
36:50Corey Noles:I noticed also that you had like four different agents that you called out. You have the deep research agent, you have an email agent and a code agent. So for people who maybe aren't developers, they can use one of those other agents as their personal assistant like you just recommended. And it's fairly straightforward for them to set that up. Could you maybe walk us through that process? Yeah, yeah, yeah. So all those agents are kind of like independent agents. You can pick one agent, talk to it. For example, the reason why we designed this email agent is we want it fully local, right? Because it doesn't make sense to design a hybrid email agent.
37:39They're like, what goes to it? Right? So that is purely local. But you can hook it up with, right now, we use it for the first one. We enable the Gmail, and you connect it and talk to it, ask your email highlights or urgent items, things like that. So give you another example. for example this um you know hybrid deep research agent um that is another agent right again you can click to it and directly talk to it and so this agent why we designed it is really i think it solves a big big problem big challenge for a lot of people you know like we love you know clubby in proper plastic and chat GPT and to do this deep research for us.
38:30You know, you have a topic and you want to find, you know, what's the best AI video editing tools and how to sequence them. What's the best price? How do I pair them together? Right? Probably that's what you want.
38:44Corey Noles:That's right. Right? And you want them to do deep, deep research, not just the search for a second to give you an answer. You want all the comparisons and all that. And you want a very indexed report. But sometimes, right, you know, this report needs to be generated in a way by combining both local and cloud information together. Because, see, if you are a CFO and you want to look at, okay, the next quarter or next year market trajectory and your own company's data How do we best position? How do we spend? Where do we invest? You know, inevitably, you need AI to look at your own data, which you cannot let a cloud AI to look at.
39:34Corey Noles:Yeah. Right? Because even if you have a data policy with OpenAI or whatever, there's no guarantee. There's somebody who reviews it for safety. That guy could potentially trade on that information. And if you're a public company, you know, the cat's out of the bag. I'm not saying that's a hypothetical scenario. I'm not saying that actually happens, but you know what I mean? There's risks involved with doing this. There's a lot of what-ifs, yeah. Yes. You know, like any CIO won't allow that, right? No. So basically, I don't know about your policy, but for me right now, I couldn't upload anything to those cloud AI engines.
40:16It was that way for a while for us, too. But over time, we've...
40:21Corey Noles:We've softened our stance. We've given in. We've given over to the machines. We can't. It's just a strict block. You don't even try. You cannot. Yeah. But we still want those time of in-depth report. How do we do it? It's a big question, right? And then we carefully designed a process, like I just mentioned, and designed a way to doing this confidential communication between local and cloud. You know, like this is the high schooler and this is like the seasoned, you know, worker. And then without telling, you know, all the data, sending all the data to the supervisor and the supervisor can draft out a plan.
41:07Okay, like for this report, and how many sections do we need, and how much information do we need to gather, and how much information we can find from cloud and web search and all that, how much information we needed to find from local, and we design schemas and all that so that, you know, by iteration, cloud model will advise the smaller models and to fulfill the tasks. So it's a pretty clever design. We make sure there's no data leakage, but we still deliver up to 90 % of the quality compared to cloud-only frontier model. You think about the best suppliers in the whole world who can do this deep research.
41:58I think the number comparison was for Office QA I think the frontier, the best solution, best in class in industry is 76%. And the hours is 68 % or 67%, something like that. Right.
42:15Corey Noles:That's awesome. Yeah. I think those are like for the individual agent level. And then if you don't want to like specifically pick any agent to directly talk to and you want default agent to route your different question. And you could directly talk to the default agent and it has this agent operator behind it. So it will dispatch the tasks to different agents to do web search and to do emails or depraise. So that's another way of utilizing using. What I like about this is that you're basically creating systems to solve problems that you actually have. Like you're like, hey, my company policy is I can't share any data to these cloud models.
43:05Corey Noles:So how can I create a deep research agent that still uses my data that I want to reference, but then complies with my policy? It's awesome. I love it. Yeah. Yeah. Interesting. Interesting. Something I'm curious about is when you look at benchmark results talking about how local processing can reduce cloud token consumption overall, at least for specific workloads, what should an enterprise or even a small business for that matter be measuring to determine whether hybrid AI is saving them money or just shifting costs into hardware and manufacturing? management instead of tokens? Yeah, that's a really good question because for cloud, you know, low maintenance.
43:59Use it, right? If you don't use it, stop it. But then for the investment of the local and you have to invest on the hardware devices and that's a one-time cost. But it's like everything else. This is like an overall trade-off in life, in many aspects.
44:22Corey Noles:You either pay up front or you pay continuously. In which bucket do I put my money? Yeah, yeah, yeah. But like we talk about, if the trends hold true, and in the very near term, and individual household and companies, big, small, and can actually comfortably host a, we call it intelligence hub, right? And that is very fundamental transformation of everything we do today. And we'll have all kinds of AI appliance and devices running around your household. But you need an intelligence hub. To become the gateway, you cannot like, you know, no matter anything, you just go to call. It doesn't make sense, right?
45:16You have to know that gateway solutions, direct the traffic, escalate it to cloud or, you know, when needed. And the similar things to, you know, like enterprises and companies. and you know it becomes I believe the future will be not like one general agent can serve all and look at our app store how many different you know thin slices thin slices vertical applications we were talking about like weather forecasting. I know we all have that. But if you live in Florida, you need a specific hurricane forecasting.
46:05Corey Noles:Or California, where I live, right? You need wildfire, you know? We need watch duty, yeah. Exactly, and today we still see a lot of this, you know, enterprise users and all that haven't fully, fully, you know, adopt AI and then bake all that into their day-to-day workflow. The reason is it's not specific enough. And like you guys are doing video editing. There's no such a thing as you like go do it, agent. No, you have to hack, right? Like manually piece together many, many, many different AI tools and tweak and you spend a ton of time. I play with it. I know the pain. So imagine there will be a lot of this type of vertical agent.
46:58And it truly, you know, you tell it what to do and you just need to, you know, more, you know, give the instruction around the taste, right? Around the, you know, how it looks, feels and all that. And instead of like manually try to hook things together. so I think there will be a booming stage of such you know very much refined polished vertical agents application coming to the world so and then yeah with that then you think about like one department this you know they may have 10 people they do one thing and why do they need And they're like a humongous general AI to serve them. They just need that one specific vertical agent.
47:48And that can be capsulized mostly in devices and serve for 10 people and another department doing something else. And that's a different thing. And most of those don't even need a state of the art model. Like the truth is the vast majority of work today can be done with much, much smaller models. Like when you think of it in terms of breaking things down into tasks and looking at the different parts, like you could do – smarter resource allocation could save so much money and so much rate limit for people.
48:24Corey Noles:I will say like the tricky part for me personally is you just don't want the agent that you use to fail at the task. You don't want hallucinations. You don't want it to be too stupid to get it done. So that's why we default to the higher intelligence model today. But once you can trust that the smaller models are, you know, intelligent equivalent, once we solve some of these edge cases, then yeah, you're right. There's going to be an explosion of vertical models, I think. Yeah. Yeah. And, you know, the other day I met a brilliant professor and now he's a startup founder. So his focus is all about cognitive learning, you know, about how to build intelligence.
49:10So our conversation is around just like what you said, like today, actually, look at all this frontier model. The intelligence level is pretty high, right? Yeah. It's really high. And has been for several generations even. Yeah. But then why can't we put it to work? Serious work. Yeah. And so I'm thinking about analogy. And let's see if you agree. I feel like it's like we have college students, graduate. But they haven't been through all these serious trainings for different specific domains. domains, they have no domain knowledge and no, you know, no hands on practice and no, you know, specific know hows and all that.
50:02So that's the last mile. And how do we train college students into experienced the worker and fast? And that's the last step of put agents to real work. And that's the domain specific. And how do we get there? and a continuous learning.
50:24Corey Noles:So that totally everything that you said, I think is 100 % true. And I think we've seen with the new models that come out that they can only focus on so much, right? You can't specialize them too much in one area because of the architecture. It will have this thing called catastrophic forgetting where it will forget the good things that it learned early on in its training if it specializes too far in any one direction. So it's sort of like that, what is it, a radar chart where it like, you know, you can have all of these like lines going off towards different specifications in a circle, but you can't go too far in any one direction or it'll like lose some of the skills from another direction.
51:04Corey Noles:So for example, a lot of them are really good at coding now, but they're not as good at writing. And we think they haven't been good at writing the frontier level ones for like a little while ever since they've been focused on coding. And everyone's like, what's this about? And, you know, I think it has to do with this. Like you you go too far one direction, you can't, you know, you can't specialize enough to be specific. So, yeah, and it has to also back into like this domain specific experience because they have specific, you know, practice and workflows and, you know, all those practices. And like you said, how to make sure agents do not lie and really deliver to the quality and to the promise.
51:47I think all of the trainings has to happen in the next one, two, three years because that determines, in my opinion, whether AI will go big or it will go the other way around. Because the big promise of AI is enhance, gratefully enhance human's productivity. And in simple terms, they have to work. Do real work. Right? Do real work. So I think we are in this mission to make these devices, make the building platforms available so that we can have more developers can utilize such devices, platforms, and to build this vertical agent and truly make AI useful, reliable, right? And carry on a lot of repetitive work, you know, for different industry verticals and fundamentally improving the productivity.
53:03Corey Noles:Doing it sustainably, you know, not burning. Sustainably. Yeah, not burning. all these tokens over the cloud. Yes, yes, yes. Well, Dr. Zude, thank you so much for joining us today. It's been a lot of fun. Thank you. Thank you so much. You guys, it's just so easy to talk to you. And I feel like it's just we are having a coffee chat. It's a lot of fun. And especially you guys are very much hands-on AI. So we are like a click. You know, just like that, right? Well, you'll have to come hang out more often. We'll have to do it. Yeah, yeah, yeah. Where can viewers go to try SuperClaw again and check out your Intel hybrid work?
53:53AIbuilder.intel.com. Okay. Perfect. Excellent. Excellent. Well, thanks again. If you liked today's video, please take a moment to like and subscribe. Don't forget to check out the Neuron's other projects, including our daily newsletter, read by more than 700 ,000 people just like you. Also, the Neuron Academy, where you can learn all about AI and how to use it in your work and life, as well as our newest sister newsletter, Robotics Insider. And that's all we have for today. Thanks for joining us. Farewell for now, humans.
From the publisher
What happens when the most capable AI lives in the cloud, but the work you want it to do is too private, expensive, or repetitive to send there every time?
Dr. Olena Zhu, Head of AI Solutions & Ecosystem for Intel’s Client Computing Group, joins Corey Noles and Grant Harvey to explain Intel’s vision for hybrid AI: systems that route each task to the right place, whether that is a local model on a PC, a larger model on an edge server, or a frontier model in the cloud.
The conversation explores why cloud-only AI may be difficult to scale, how an orchestration layer decides where work should run, and how Intel’s new SuperClaw beta combines local and remote intelligence. Dr. Zhu also explains how hybrid agents could conduct deep research without exposing confidential data, why routing and auditability matter when agents fail, and why the next wave of useful AI may come from specialized agents built for specific jobs.
Listen for a practical look at the tradeoffs among capability, privacy, cost, reliability, and control, and what it could mean when powerful AI agents begin running partly on the computers we already own.
Try Intel SuperClaw: https://aibuilder.intel.com/#/superclaw
Intel AI Builder: https://aibuilder.intel.com/
Intel SuperClaw GitHub: https://github.com/intel/intel-ai-builder/tree/main/superclaw
Intel’s SuperClaw overview: https://newsroom.intel.com/opinion/solving-the-agentic-ai-trilemma-cost-scale-and-data-security
Subscribe to The Neuron for daily AI news and analysis built for humans: https://www.theneurondaily.com/
Sponsored by Nerdio: Get a demo of Nerdio Manager for Enterprise today:
