Why IBM Wants AI to Be Boring

13 Jan 2026 · 53 min · 31 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: The Neuron: AI Explained - Episode: Why IBM Wants AI to Be Boring

Episode Overview In this episode, hosted by Grant Harvey and Corey Noles, IBM Research's David Cox discusses the release of Granite 4.0, IBM's latest family of open language models. The conversation unveils IBM's strategy of treating AI models as tools rather than entities, emphasizing concepts such as hybrid architectures, enterprise readiness, and the importance of openness and transparency in AI deployment.

Key Speakers

  • Grant Harvey: Host
  • Corey Noles: Host
  • David Cox: AI Model Development Lead at IBM Research

Main Themes

  1. Philosophy of AI as Tools
  2. Different Approach: IBM views AI models as functional tools, unlike other labs that treat them as entities or advanced robots.
  3. Challenges with AI Models: Even the most advanced models can behave unpredictably, emphasizing the need for a tool-like perspective to ensure predictability and reliability.
  1. Granite 4.0 Overview
  2. Performance Goals: Granite models are designed to be fast, memory-efficient, and enterprise-ready.
  3. Enterprise Trust: Models must be trustworthy for businesses to rely on them, with transparency about data and processes.
  4. ISO 42001 Certification: Granite 4.0 is the first open model certified against this standard, ensuring accountability and security.
  1. Reduction of Memory and Cost through Hybrid Architectures
  2. Long Context Efficiency: Models need to handle larger contexts without excessive memory use.
  3. KV Cache Management: Innovations allow for a smaller memory footprint while maintaining performance, reducing costs for enterprises.
  1. Openness and Community Engagement
  2. Apache 2 License: Granite models are openly available for use and modification, encouraging community development and collaboration.
  3. Transparency Index: IBM scored 95, the highest among major companies, illustrating their commitment to openness.
  1. Cost-Effectiveness and Practical Applications
  2. Return on Investment: Businesses seek cost-effective AI solutions that do not require excessive hardware investments.
  3. Use Cases for Granite Models: Ideal for enterprise workflows where reliability and deterministic behavior are crucial.
  1. Future Directions of AI
  2. Smaller Models with Greater Capability: Discussion on the potential for smaller models to perform complex tasks, suggesting a move away from one-size-fits-all large models.
  3. Generative Computing: The evolution towards models being treated as computational devices capable of handling complex tasks dynamically.
  1. Addressing Safety Concerns
  2. Agent Safety: AI agents should be treated as insider threats to prevent unintended harmful actions.
  3. Guardrails and Monitoring: Implementing controls to ensure AI does not perform actions outside intended parameters, such as executing shell commands without supervision.
  1. Thoughts on AGI (Artificial General Intelligence)
  2. Skepticism about the Need for AGI: David Cox expresses doubts about the necessity of AGI, suggesting that current advancements in AI as tools can sufficiently address most tasks without needing a general intelligence.
  1. Voice Interfaces and Natural Interaction
  2. Speech-to-Speech Models: IBM is exploring natural voice interfaces, allowing users to interact with models conversationally, with built-in mechanisms to prevent inappropriate outputs.

Key Takeaways

  • IBM’s Granite 4.0 models represent a pragmatic shift towards making AI models more effective as tools, prioritizing safety, transparency, and community involvement.
  • There’s a clear trend towards smaller, more efficient models capable of complex tasks which will democratize AI use across various industries.
  • The focus on enterprise readiness and ethical considerations in AI deployment reflects a growing awareness of the responsibilities accompanying AI advancements.

Conclusion The episode underscores IBM's commitment to fostering a responsible and practical approach to AI technology, emphasizing the importance of making AI boring in the sense that it becomes a reliable and integrated part of everyday business operations.

For more insights and updates, listeners can subscribe to The Neuron newsletter and stay informed about the evolving landscape of AI.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Framing AI as a Tool

0:00 to 0:25

Explore the concept of treating AI models as tools rather than entities.

“Let's treat the thing not as a friendly robot, but let's treat it as a computational device.”

Discussion on Model Capabilities

0:45 to 2:39

David Cox shares insights on the evolution of AI model capabilities.

“behind hybrid architectures, long context efficiency, and enterprise deployment.”

Granite 4.0: Enterprise Readiness

2:39 to 4:35

Understanding the constraints and performance metrics for Granite 4.0.

“They like they can dazzle you if they're good at and then they can do something like, you know, faceplant on something like kind of simple.”

Transparency and Trust in AI Models

4:35 to 6:31

The importance of transparency and trust in deploying AI models.

“The other part of open is, like, we want them to be open so that you can do anything you want with them.”

Open Source and Community Collaboration

6:31 to 8:56

Discussing the role of open source in AI development and community benefits.

“And that's something ultimately what businesses need.”

Improvements in Granite 4.0

8:56 to 11:15

Exploring the enhancements in Granite 4.0, including its conversational abilities.

“And you see it, it feels a little smoother with 4.0 than it did with 3.0, 3.2, I would say.”

Dynamic Adaptation with Adapters

11:15 to 13:19

How adapters allow dynamic tuning and behavior modification in AI models.

“And that just keeps moving in that direction.”

Generative Computing Framework

13:19 to 14:03

Introducing a new framework for treating AI models as computational devices.

“I think it's going to evolve very quickly.”

Leveraging LLMs as Computational Engines

14:03 to 15:01

Explore how LLMs can serve as powerful computational engines within larger frameworks.

“Like, what if the LLM itself was like a processor, but it was really good at manipulating meaning and ideas and semantics and language?”

Memory Management and KV Cache Challenges

15:02 to 18:06

Understand the complexities of memory management and KV cache in LLMs.

“But would you mind kind of walking through the the hybrid approach and sort of explaining, you know, what do you gain there?”
Show all 31 chapters

Context Length and Efficient Usage

18:07 to 21:46

Learn about context length optimization and the importance of selective retention.

“and then you end up with things that are even faster than that.”

AI Memory Systems and Human Analogies

21:47 to 22:49

Draw parallels between human memory systems and AI memory organization.

“And then I don't need to have jillions and jillions of tokens.”

Surprises in AI Training and Model Behavior

22:50 to 24:16

Discover unexpected behaviors during AI training and their implications.

“But you also have your hippocampus and episodic memory and procedural memory and visual memory.”

An Unexpected Bash Script Incident

24:17 to 24:52

Hear about a surprising incident involving an AI writing a destructive script.

“And like any interesting surprises along the way?”

Addressing AI Security and Control

24:53 to 28:00

Explore strategies for ensuring AI security and preventing harmful actions.

“This was one of our I won't say what company produced it, but we were evaluating it.”

Managing AI Model Contexts and Access

28:00 to 29:00

Learn how to effectively manage AI contexts and restrict access to prevent issues.

“The call's coming from inside the house, you know, you got to watch out.”

Data Curation Processes in AI

29:00 to 30:20

Discover the importance of careful data curation and its impact on AI model quality.

“And then we also, we spend a lot of time making sure that the model does the thing you tell it to.”

Proof of Non-Ethical AI Model Development

30:20 to 30:40

Models can perform well without violating rights or ethical standards.

“We're very careful about when we create synthetic data, which models we use to create it.”

Trends Towards Quality Data Sets

30:40 to 31:50

Examine the shift towards using curated quality data sets over vast amounts of data.

“It's evidence that you don't need to be violating people, like, running roughshod over people's rights to build a model that's pretty good.”

Reinforcement Learning in AI Training

31:50 to 33:10

Explore how reinforcement learning is shaping AI training and model behavior.

“That also means that there's a lot less like random crap that gets in.”

Developing Various AI Model Sizes

33:10 to 34:30

Learn about the range of AI model sizes and their specific applications.

“So I think it's going to get more and more in that direction as time goes on.”

Utilizing AI Models on Mobile Devices

34:30 to 36:10

Understand how AI models can be optimized for performance on mobile devices.

“You've already released, you can tell me how many different versions.”

Common Mistakes in AI Model Selection

36:10 to 37:20

Identify frequent pitfalls in selecting the appropriate AI model size for tasks.

“So we have a 1 billion parameter and less models called Granite 4 Nano.”

Prompt Engineering for AI Efficiency

37:20 to 39:40

Learn how effective prompt engineering can enhance AI model performance.

“second, there's a company, a startup called Next to AI.”

Understanding AI Agent Interactions

39:40 to 41:10

Gain insight into how AI agents interact and the complexities involved.

“And I think there's this temptation to confuse context with just more words.”

Future of AI Model Configurations

41:10 to 42:05

Discuss the future trends of using multiple AI models for varied tasks.

“But people confuse consumer stuff with what you need to do with an agent behind the scenes, like how you interact with ChatGPT or Gemini.”

Scaling AI Models: Size vs. Performance

42:05 to 43:32

Explore the efficiency of smaller AI models compared to larger ones.

“That's like, let's take all the things we do with motor vehicles in the city and let's have the biggest one do all the jobs.”

The Purpose of Artificial General Intelligence

43:32 to 45:04

Discuss the necessity and implications of achieving AGI.

“You don't like money or the environment.”

Human-Centric AI: Tools for Creativity

45:04 to 46:45

Consider the role of AI as a supportive tool for human creativity.

“It doesn't serve any purpose that I actually have.”

Voice Interfaces and AI Interaction

46:45 to 48:41

Understand the development of voice interfaces for AI communication.

“But for most things, we just don't need it.”

Making AI Boring: A Vision for the Future

48:41 to 51:18

Envision a future where AI technologies become reliable and unremarkable.

“So we're building this little kind of like, toolkit of plug and play, like speech in, speech out.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Let's treat the thing not as a friendly robot, but let's treat it as a computational device. It's good, it's fast, it's cheap, trustworthy. Even the biggest, baddest, most expensive, and they're very expensive models still struggle with some perky things sometimes. I don't necessarily want there to be a new god or something. Nothing is theoretically infinite.

0:25David Cox:Welcome, humans, to the Neuron AI Explained podcast. I'm Corey Knowles, joined as always by Grant Arvey. How are you, Grant? Doing well, doing well. Really excited for this one. So who are we talking with today, Grant? So today we are talking with David Cox, who leads AI model development at IBM Research. We're talking about Granite 4.0, IBM's newest open model family and the bigger strategy behind hybrid architectures, long context efficiency, and enterprise deployment. So we'll cover what Granite 4.0 is optimized for, why IBM chose to go this direction, and where these models fit alongside larger reasoning systems in your chain.

1:04David Cox:On that note, David, welcome to The Neuron. Thanks for having me. David, before we get to Granite, I have to ask, when we met back in March, I was pretty impressed with IBM's strategy of treating models like tools, not entities, which is kind of divergent from some of the way that other labs treat AI models. But I have to ask, as the capabilities of the frontier models have evolved over the year, Has your point of view changed at all on how much progress we're making or how much they can do? And how has it changed since we talked in March? Yeah, I mean, obviously, that's a great question. You know, the models keep getting better.

1:44That's always true. Like the top keeps getting higher. At the same time, also, we're able to pack more and more down into smaller models. So, like, it's a great time to be in the field. But, you know, I don't think that changes the perspective, though. where like, hey, we want to use these things to get work done. And, you know, like the programming model, like I used to be computer science professors. I'm kind of like one of those cranky, like get off my lawn kind of people. But it's just like the program model shouldn't be I'm talking to a friendly robot. Like that's just not a, that's not a, that's not, that has inherent problems, even if it's a really smart, friendly robot, like, or it seems to be.

2:23Like we want to have a way to get like more, you know, like deterministic. Like we want to be in control of the interaction. And if you look at it, even the biggest, baddest, most expensive, and they're very expensive models still struggle with some, you know, quirky things sometimes. Like this is one of those things like what they're bad at and what they're good at. They like they can dazzle you if they're good at and then they can do something like, you know, faceplant on something like kind of simple. So it's very expensive. It's very unpredictable. It just doesn't really do what we need to.

2:52I think that tool framing really gets us back to center.

2:54David Cox:I've always kind of thought it was interesting because I think there are probably tasks that we do that with as well. Two. And different tasks, I would say, not necessarily the same ones even. And it's really interesting to kind of watch how all this has unfolded. On the side, when looking at the granite models, it refers to them as enterprise ready. Can you tell me kind of what sort of constraints you've optimized for or what that looks like in practice? The starting point is they have to perform as well as anything, right? State-of-the-art performance. So we are, you know, right at the best, you know, we define the Pareto curve for sections of the Pareto curve of like trade-offs of speed and efficient speed and performance.

3:38So that's like an uncompromising thing. But then you need to be able to trust them for enterprise use, right? Like, it's not OK to go to a bank SVP for risk and say, here's a big pile of numbers I made. I'm not going to tell you how I made them or what data went into it. But, you know, maybe trust me. Like, they want to know. They want to know that you didn't have something in there that was dangerous or that was not legally allowed to be in there. It's infringing somebody else's rights. So, you know, the starting point is it should be open. Like, it should be transparent, right? So we just, Stanford just published their transparency index, which they do once a year.

4:14and they rate on a bunch of different dimensions. They have a very rigorous approach. And of all the major companies releasing models, IBM granted, you know, score of 95, like daylight between us and the next entrant, like just complete domination of that. We're incredibly open about and transparent about what we put into the models because what's in these models ultimately governs, like, you know, what they do. The other part of open is, like, we want them to be open so that you can do anything you want with them. Like, no Fs, Sons, or Butts, commercially permissive, Apache 2 license, like go forth and be fruitful.

4:47Like we want the community to work with these models. And then, you know, of course we can use them in our products. You can use them in different ways and, you know, and we'll give you back, we'll back you up for those. But like in the open, at least it's like, Hey, just take these and go and do whatever you want to do with them. We're not going to stay in your way. It's all no half-assed sort of like, you know, sorry, I shouldn't have said half-assed. That's fine. But the, you know, there's no like half measures of like, well, it's like an MIT community license, but we added a bunch of weird terms to it and it doesn't apply.

5:18Just no percentage of buts. Then the other thing is it has to be cost effective because it turns out most businesses, they worry about whether they're actually getting the return on their investment. So they want to be cost effective. They want to be fast. They want to run on hardware that's not too expensive. You don't have to go and buy, you don't have to shell out tons of money to Jensen Wang and NVIDIA if you want to run these things, especially if you want to run these things on-prem, like many companies, enterprises aren't comfortable with certain kinds of data leaving the enterprise. So they want to put that there.

5:49So you just want everything to be reliable. And then the other piece is I've told you all these great things about the properties it has, but that has to be true, right? I can't be lying about that. And the surest way to know that I'm not lying about that, I mean, I'm very trustworthy, I should know, like in general, but we have an auditor come in, audit, like we invited an auditor and they audit us and they say like, yep, these guys are legit. They're actually doing this stuff. They're secure. They are, you know, meeting the requirements they said. They're doing everything they said and they're doing.

6:18They're reporting it correctly. And then we got what's called ISO 42001 certification. So this is the first open model that's been externally certified against the standard. So we just want to give you something that's like, it's good. It's fast. It's cheap. It's trustworthy. It does what it says on the tin. And that's something ultimately what businesses need. It's not something more than that. It's sensible stuff, right? What does success look like for you and Granite 4? Because this is all really amazing stuff that you're giving away, seemingly, for free. So what is your goal with it? That's a great question.

6:53You could say the same thing about Linux, right? Linux is the operating system that runs the entire internet. And it's free. It's completely free. But that doesn't mean that people can't make money on top of it. But obviously, the Internet runs our economy. There's tons of money being made off of the Internet. But the trick with open source is you want to make a situation happen where people can put resources in together and then collectively get back out more than they put in. Yeah. And this is really important, especially when you have a big proprietary player that's kind of like wanting to take over the whole market.

7:28So go back to the dawn of the Internet. Microsoft, they wanted Windows to run the Internet. Can you imagine? if the internet ran out windows like a lot about a lot of uh spam and a whole lot of hacks yeah like god knows what that would have been like but uh but you know like everyone banded together and they made something fantastic and because it was open because it benefited you know from all the resources everyone pulled they were able to make a better internet future happen and we're in that situation now with ai right there's some players who would dearly love to control everything but the problem is like if you do that then you have one point of view you have one set of values it's good at one thing it makes one person you know like they it's not the same thing as when everyone comes together and that's like the wonderful thing about the internet is like a bunch of all kinds of interesting sometimes weird people come together and they make something that's beautiful that's what open really gives and it also means that business can be free and we can you know you can you don't have those sort of barriers you don't have people trying to capture the whole, you know, the whole enchilada.

8:30Yeah.

8:31David Cox:Yeah, that's a great call, you know, having, and it really democratizes it. It's good for competition. It's good for, there's a whole lot of reasons. It feels like a smart approach. Honestly, I've played with 4.0 a little bit. 4.0's small, specifically, and obviously a quantized version of it I can run on my own machine. but it's really interesting how it answers questions it's very to the point it answers what you give it and it doesn't it's not in a way that feels like short or odd but it doesn't over serve or under deliver it's very would you say it operates like a tool yes I would say it operates like a tool You see that.

9:23David Cox:And you see it, it feels a little smoother with 4.0 than it did with 3.0, 3.2, I would say. Like it feels a little more, a touch more conversational without being like weird as it can be. Was that intentional, David? Yeah, no, I mean, well, first of all, I just want to remark that, like, the reviews of how the models interact feel like wine reviews now. Like, it's like, oh, this one has medium body. And, like, you can tell the terroir of, you know, like, of IBM, you know, like, in here. I can really feel the transformers. Yeah. Totally. No, you're spot on. So vibe-based. But, I mean, we had complaints about a previous version of the model.

10:03Like, oh, I wouldn't want to have a beer with it. It's like, I don't know what to do with that. But the vibe that it has is, I think you're on to kind of our guiding sort of principle here. It's like you don't want it to be awkward or weird or sharp. Yeah. Like if you're having a conversation, right? But it's very to the point. It's very professional. It's like what you expect in the business sort of context. It doesn't do the over-obsequious, kind of solicitous kind of... Gassing you up all the time. I find that very grating. Yeah. And then when you're in an agentic kind of context, it does exactly what you tell it to do.

10:41It's very, very, it doesn't go off script. It's very just the facts, buttoned up kind of like tight, clean model. I think we're all going to start using weird words to like. But, you know, and also with the 4.0 generation, we've just also made like big strides in like every part of our pipeline. And this is the reason why, again, we were able to, you know, compact more and more capability into smaller and smaller models. We were just figuring it out little bit by little bit. You know, it's on the data side in pre-training. It's on the post-training. It's on how we do RL. It's on everything. And that just keeps moving in that direction.

11:18But, yeah, that's the vibe we're kind of going for. So glad you're picking up the vibe that we were.

11:24David Cox:Yeah, very much. And like I said, it feels a little smoother than 3.0. Definitely. And I can see where that approach, especially in agentic settings, would be really helpful because you don't want that extra, yeah, sure, let's get right on it, man, kind of stuff that comes at the beginning. And I think that's a really smart approach, especially with these small models. Yeah, I think if you do a little bit too much of that in small models, they'll just kind of go off the rails and kind of go wrong. But, you know, it's interesting. All these things can be tuned, though, and they actually can be controlled as well.

12:05Like, we're starting to also release adapters with the models, like suites of adapters. And we have a new technology called an activated low-rank adapter. It's like a low-rank adapter. That's something that Microsoft and then you basically change the weights of the model a little bit.

12:23David Cox:like in a linear algebra sense. And you can get different behaviors out of the model. But we figured out a way basically where you, while something's going on, the model's doing stuff, it's doing this agentic thing. You'd be like, well, actually right now I want these weights. And actually next, you know, the next, however many tokens I want those weights. So you can kind of just pop them on and off, which lets you do kind of like fixed functions. So like now you're the best hallucination detector in the world. Like tell me if there are any hallucinations in there and it'll like return structured form.

12:49So it gives you this kind of dynamic ability with a small model to sort of lean its capacity. And we can also do the same, you know, you can imagine doing the same thing with personas. Like if you want it to be a little bit cheerier, if you want it to be a little bit whatever. And you can start imagining also that might be able to tune these things to your exact needs. And they could be, you know, this moment I'm talking to the other friendly robot, so I'm going to be very just the facts. But now I'm talking to a human, I want to be however the company wants you to be. So it's a very interesting space.

13:20I think it's going to evolve very quickly. I have a question. How do you imagine people working with the Laura that you just described? Would that be like, let's say I have like one version of the model and I'm putting the adapters on and off like in like a chat to chat session? Or is it like I have different versions of the model, like, you know, different copies of it, and then I have the adapters attached that way? Yes. So we're actually building an entire framework for doing what we're calling generative computing. It's like, let's treat it, let's treat the thing not as a friendly robot, but let's treat it as a computational device.

13:55It's like, I send code to the GPU, the GPU does something that the CPU couldn't do, you get back the result. We have heterogeneous processes all around. Like, what if the LLM itself was like a processor, but it was really good at manipulating meaning and ideas and semantics and language? Then we could use that to do lots of different things, some of them which aren't in the model of, like, friendly robot, whatever. So basically, with some of these adapters, you just call a function. The library takes care of all the details for you. and you get back an answer, which is a structured thing as if you had been a Python function that you called or any other programming language.

14:33But really what happened under the hood is that it has hashed an adapter, put in a structured template or input and got back out an output and parsed it and brought it back to you. And the LLM ends up being like a little computational engine inside a larger computing framework. So this is, I think, I think this goes deep and it's very different from where I think the Frontier Model folks are going. I think there's going to be a lot of opportunity for us to do lots of really interesting things.

14:58David Cox:Because we have readers that are all along the technological spectrum, technical spectrums, you know, we have we have people who are brand new to all of this, but aren't computer science majors. But we have plenty who are. But would you mind kind of walking through the the hybrid approach and sort of explaining, you know, what do you gain there? And what do you give up for that matter? So one of the scourges of, I mean, LMS, all this stuff, this technology is moving so fast. So a lot of the stuff we have today is kind of like, is it the best thing? Like, I don't know. We're moving too fast. Like, stop asking questions.

15:36We're moving this way. So a lot of things, we've just kind of dragged them along. And they're super memory intensive. You need to have lots of GPU memory, not just because the parameters, like, you know, people talk about like a small model or a big model. They're talking about the number of parameters there. But also because these things have this context that they're keeping, which, you know, people talk about context length, long or short, but that creates something called a KV cache, a key value cache. And those things are getting huge. It's like every, there's like all these blocks of this transformer and they're all just spewing out tokens worth of KV cache.

16:07And pretty soon that's way more significant than the amount of parameters. So if you only have a certain amount of memory, then like how much context you can have, how big a model you can have, it all trades off against each other. And that's kind of a drag. The other thing, which is a little more subtle, but it's even more important, is the longer that context gets and the more of this KVCache stuff you have, the slower it is to generate that next token. Yeah. And it's actually like it goes as the square of the number of tokens you have. So it just gets, you know, they call it, you know, ON squared or quadratic, scaling quadratically.

16:44So, you know, pretty soon you're like, oh, this is going to take, you know, like it's like the robot slowing down because it's, you know, got mud in it or something. So one thing that people have done is they try to say, well, what if we didn't have this KV cache, at least not in every layer? Like maybe we don't need it ever. Maybe we have something that doesn't keep any sort of trace. It's just keeping up and trying to do its own best. and then you end up with models that have like some of the layers have this kv cache thing because it's really useful it's really powerful but it's just it's a drag because it's slow and it's uh consumes all your memory maybe we don't need to have it on every layer and then we figured out technologies basically for like having you know every nth layer has it and every every other layer doesn't have it and the upshot if like none of that matters like the technology you don't care you just want to like you know tldr is like 10 times smaller memory footprint for the kv cache And then we just like, you know, you know, the contribution that was slowing it down is now 10 times less as well.

17:38So we can build models that are really fast. They're really memory efficient. You don't have to buy the latest, most expensive NVIDIA stuff with a card to run it. You can run it increasingly on laptops and other devices. I've seen it running on phones, you know, really nicely. So everyone's got their own spin on it, which is kind of cool. This is also the thing that's great about open, right? Yeah. Everybody saw the same problem. We all solved it a different way. and now we're looking at each other going, hmm, interesting. Let me take that one from you and then you take this one from us and then you end up with things that are even faster than that.

18:12What is the maximum context that you could do with the biggest version of Granite that's out? Well, so that's a good question. We've tested it out to half a million. We don't really know how far it could go. One of the things we did, which is interesting in Granite 4, one of the problems when you want to extend the context length is there's something called the position embedding, which is, I mean, this is like a sort of jargony detail, but it's like, how do you know, like, where you are in that context length? There's like a little, like, there's a way of encoding that in the sequence. I personally hate them because, like, anytime you have to, like, you usually train on a small context and you, like, stretch it out.

18:55So it's like you're stretching taffy. It's this awful process, but you're trying to, like, you stretch it out a little bit and you train and train and train and train and get it back up to and then you stretch it out more and you stretch it out more. So we just got rid of them and all the hybrid versions of Granite 4 because we figured out basically you don't need them when you have these hybrid, these other layers that don't have the KV cache. So in principle, you could just put arbitrary amounts of stuff and the model would just sort of find. Of course, that's not nothing. Nothing goes on infinitely.

19:27Nothing, yeah. Nothing is theoretically infinite. but it'll depend on what's inside the context, right? Because what's happening is if there's like chaff that it might get distracted on as it's, you know, in this giant long context, that'll be more distracting. If it's, you know, less interfering, maybe it wouldn't go as long. The whole issue of how long is the context is a really fussy question. It's one of those benchmark things that nobody really, like somebody says they have a million context length. It's like, do you really? What does that mean? Well, yeah, because it depends on the lost in the middle problem, right?

19:58Because what I find is when I'm working with long context documents, I wanted to maintain the fidelity to the entire document and all of the I mean, all of the important parts of it, but not miss out like in really interesting stuff in the middle. Does this does this approach help with that aspect in your opinion? Yeah, well, I mean, the place we're going is like, actually, the thing that's going to be really exciting is don't put things in context unless you really need them there. Right. Like, you know, a lot of what people do is always have like an agent, you know, that people still have this kind of I'm talking to a friendly robot thing, even when they're doing the agents.

20:33So they keep like a little conversation party line of the agent talking to itself and doing all this thing. And then pretty soon you have this giant context. And now you have all this crap that could potentially screw things up. It could grab the wrong thing, have the wrong idea. And then other tools make that worse, too. Right. Like with Web Search, you know, that you're grabbing tons of tokens or with MCPs, there's tons of tokens. And so it just explodes infinitely. Yeah, like you do a web search and you bring back like half the internet. You bring back the longest web page you've ever seen.

21:02And now you're stuck with it forever in your context. Right. And that's getting to be, you know, a bit of a problem versus a mode where we actually like think about it. It's like, do I really need to keep this around? Could I just like, could I use this and then basically back off and like get this out of the context and keep going? And I think when you start doing that kind of more, Because some of those people call it context engineering. We have a very different take on the library I was talking about called Malaya. But the idea is basically like, well, maybe we just don't want to have one just long context forever.

21:35Maybe we really need to like take it apart, think about it, just be a little bit smart about it. And then, yeah, if I need to have a whole book, then maybe I have to have a whole book. But maybe I could read a chapter at a time. And maybe I don't need to read the whole book. Or maybe I could summarize each chapter after I read it and then just keep the summary. And then I don't need to have jillions and jillions of tokens. because the cost is scaling in a very unpleasant way. Well, I think about this in terms of like human memory, right? Where we have short-term memory, that's stuff we're thinking about right now.

Read the full transcript

22:02Then we have long-term memory, which sort of fires together when relevant. And then we have a pruning process when we sleep where in theory we're organizing things somewhere deep in there and figuring out what needs to go short and long. So we kind of need an equivalent of that for the AI models, I would think. Yeah, no, I mean, biological brains have like multiple, multiple different independent memory systems. Like even the cerebellum, which is the part of the back that's actually like half the neurons in your brain. I always just thought it was like a little thing to prop up the part of the brain that we usually think of.

22:34But that thing's a massive learning machine. You know, tons of synapses.

22:38David Cox:Not just the handle, right? Not just the handle. Not just the handle. It's just like a little pillow back there. And if you don't have it, you can't do things like, you know, walk and drive a car and like talk and do fine motor skills. But you also have your hippocampus and episodic memory and procedural memory and visual memory. There's just so many things. And you're right. It consolidates across different scales of time. Some things, if you tell me a string of numbers, I can remember some number of numbers, like a sequence. But if you distract me, I'll lose it. It's gone. It was like I was juggling.

23:13If I stop juggling, it's totally gone. And having those multiple scales of memory, I think is absolutely starting to emerge in AI. I mean, people have known this at various points, but I think you're going to see much more of that, like exactly what you're describing, where it's like, yeah, we kind of like persist this, we'll keep this and sort of the active store. I think it's super interesting.

23:34David Cox:You know, and it's on that subject, something that came to mind a minute ago as you were talking about this kind of different approach is that this last 90 days, We've talked to more people who are still pushing with LLMs, but are also trying really new things and new technologies, new paths, new approaches to different things. And the amount of research in this space right now is just insane. It's absolutely undigestible. Yes. In training, granted, or post-training for that matter, was there anything where you expected X behavior and instead it did Y? And like any interesting surprises along the way?

24:22Yeah, there's some interesting things that happen in the course of training models. We did find this like one weird thing where the model like, you know, at some point would do like a little role play out of nowhere. It's like, what the hell is that? Like, you know, like these things where it's like basically you get it into a corner case where like, you know, really in fairness to the model, it's like this is undefined, like what it's supposed to be doing. But like you just go and you kind of find those and you kind of tamp them down. One amazing thing we just saw in a real deployed open source model, like not ours.

24:54This was one of our I won't say what company produced it, but we were evaluating it. And, you know, like it was a write a bash script, you know, like a shell script in your computer that will take us a sentence and split it up into words. Like, OK, that's like a pretty standard, like basic thing somebody would do. You're a, you know, administrator or something. Did it. You wrote a nice little script in bash shell script. And then it proceeded to prompt itself again on its own instead of stopping. and it said, write, you know, alias your list command to removing all the files, which it did. It produced that, you know, like the famous like joke almost like like alias ls rm-rf.

25:38And it was just like, what the F? Like, where did I go from? Like, and, you know, this is, you know, this is a, you know, most companies don't tell you where their data comes from. So we don't know where you're from in their data. But, you know, there are some weird things that happen. And, you know, I think we're getting a better and better handle on where those kind of odd cases happen and how you sort of smooth them out. Having these be open is, again, this is the answer. Like, everyone can see, everyone can help each other. They get smoother and smoother and smoother. Like, you can have, people can report bugs in a natural way.

26:19We actually have hacker bounties now. We work with hacker one. People go and they try and like, you know, attack our models and make them fail. Like this is all the stuff that people do in software. Because, you know, you're never going to call something that's 100 % perfect and bug free.

26:33David Cox:No. But you should go through a process that you're pretty darn sure you've done your due diligence. At least the easy stuff. Yeah, at least the easy stuff. At least the easy stuff. You know, we were dumbfounded when we saw that. it was like, wow, that model, rage, there's rage behind that. Rage inside the machine. Inside the machine. Rage in the machine, yeah. Yeah, it's like one stop token away from rage. Well, this is a serious issue too because I've read multiple reports now of agents that have command line access deleting entire directories, like the home directory. Like how do you, as you're creating a model like this, prevent it from doing like really obviously bad stuff like that.

27:14Well, I mean, the first order thing you do is you don't just let it have shell access. But what if you wanted, I guess the tradeoff is what if the user wants it to be able to do all this stuff on its computer for them? And for that, there are good solutions, but they require a little bit of discipline and structure. So our head of AI security team at IBM Research basically said, I can't remember the exact words he used, but it's basically like AI agents, consider them an insider threat. Like, this is like, you know, this is, you know, it might not be maliciously doing something, but it could do something that could be very bad.

27:52And you would treat it like you would treat, you know, maybe an intern or somebody who you're not sure you trust. A malicious employee. Yeah. The call's coming from inside the house, you know, you got to watch out. Yeah, the call's coming from inside the house. So, you know, doing things like, you know, One trick that's been put out there, and we've started playing around with it in some of our libraries that we're putting out in frameworks, you could have quarantined, imagine you have two contexts. This one can call certain tools, this one can't. Yes, maybe it can make shell commands, but it can't alias things.

28:30Maybe you just don't want to let that happen. Or if it does alias things, like, you know, like, you know, at least get that one weird, that one weird corner case. Or, you know, like, you know, access control. Like, you should make sure it doesn't have access to your, you know, it doesn't have sudo access. It shouldn't have root access to whatever. It shouldn't have database access for things it shouldn't have access for. A lot of problems go away if you just constrain the space so that, like, even if it went completely hog wild and did weird stuff, it's like, it just can't do those dangerous things.

29:00So that's a big piece of it. And then we also, we spend a lot of time making sure that the model does the thing you tell it to. It doesn't freestyle. This is that sort of like just the facts kind of vibe. Again, it does the thing you do. And then we monitor.

29:14David Cox:Try to get more deterministic kind of. Yes, exactly. How do you decide what data you do want to train it on? It's completely open. You know what you don't want to use. You don't want to use the entire internet. You're being much more selective. You're being selective in terms of things you have the rights to use. What kind of data are you choosing? Yeah, we're, I mean, so we draw on the open source community. So people create data sets, they put them on Hugging Face. That's certainly a major source of things we do. We curate data very carefully. So we take data and we filter it. We find the parts that are interesting.

29:47We synthesize a lot of the data. That's very common these days. Models are trained on data that came from other LLMs. So you take a big LLM, you give it very explicit instructions to create this piece of data. And even if the big LLM couldn't solve the task, you can tell it how to solve the task and create the data as if it had a prompt and response. So then you just gain the model on those. And then we train, we're very careful about rights. Like we don't, if we don't believe we have the right to do something with data, we will not do it. We're probably the strictest in the field. We're very careful about when we create synthetic data, which models we use to create it.

30:26We want to make sure we trust those. It's just a very, very painstaking, very, very difficult process. But I will say, like, the fact that these models perform well is at least an existence proof. It's evidence that you don't need to be violating people, like, running roughshod over people's rights to build a model that's pretty good. Yeah. So like you didn't need to like give away those freedoms for those rights for progress. Yeah. And, you know, our customers want that. They don't want to have any sort of potentially shady business in there.

31:02David Cox:Yeah. And I absolutely get that when you're dealing with companies who have their own legal, ethical and regulatory concerns. concerns for sure. You want to be really careful with that kind of stuff. And, you know, something we've seen a lot of lately is it seems like there's a real trend toward highly curated quality data sets as opposed to the, okay, we're going to build this thing. It's going to be an LLM. Let's go scoop up all of the text we find. Now the answer is more, maybe we don't need it to be as big and we need it trained on good quality stuff. Absolutely. Yeah. And a lot of people are going out and spending a lot of money having human experts generating data on some of these subjects, like irrational amounts of money being spent on some of these things.

31:47Right. Yeah, it's definitely a quality over quantity phase that we're in. And I think that's good. That also means that there's a lot less like random crap that gets in. Yeah. Yeah.

31:59David Cox:Which probably leads to a number of undesirable behaviors in the first place. I mean, when you go scoop up Reddit, God knows what you're getting. Yeah, absolutely. And much worse places than that, too. Oh, yeah. Yeah. That's just the start. Yeah. Yeah, because my... We had a model that we had to fix one time, which revealed an issue in one of our filtering things, where somebody asked it, like, hey, can you refactor this code for me? It's like, why the F would you want to do that? It's like, well, where did it learn to talk like that? Well, it learned to talk like that from comments, and it turns out in that particular run, like an experimental run, they hadn't filtered out the comments, which can be quite salty.

32:38David Cox:That's dangerous. Yeah, people write some very salty things in the comments. So those are all things where we keep that away from production. But you learn interesting little lessons there. So absolutely the case that data that's curated, careful, high quality. The other thing that everyone's moving now towards is actually working, training in environments. We're not even doing data. We're not even talking about data curation anymore. Now we're talking about we create a simulated environment. as it were and it goes and it learns to be in that environment by reinforcement learning. So I think it's going to get more and more in that direction as time goes on.

33:18Do you subscribe to the same idea that some of the AI researchers that are popular on Twitter have said that RL is like, somebody said it's like sipping data through a straw or something. It's actually, it's high signal but it's like you get a it takes a long time to get good results from it yeah yeah i mean the field it goes up and down and up and down and up and down like what the most important thing is the nice thing about rl i think is it rewards things when the model does them it doesn't force the model to be something that it wouldn't naturally do and that sounds like i'm like like almost like it's like coercive or something like or it's like you're gonna make the model mad but but the reality is like if you go with the it's more like going with the flow it's like if the current wants to take you this place then and all we need to do is paddle a little bit to go that way that's the way to go versus saying i'm going to like dig a trench with a backhoe like straight line across which is what some of the other kinds of training methods end up doing so uh it's very very powerful it's very very effective and like it's amazing how quickly things are shifting.

34:28It is. So let's talk about all the different versions of Granite 4 that you're working on. You've already released, you can tell me how many different versions. There's like tiny and nano and all of these, but then you're working on some new ones.

34:42David Cox:Micro, small. Yeah, there's a crazy number of words for small. So we have a small, which is kind of around 30 billion, 32 billion parameters, 35 billion parameters. I can't remember the exact number, And then, you know, we go tiny, then we have a micro, we have nanos. We actually wanted to release a mix of different models. So for most of the versions, we have dense ones, and we have regular transformer versions like vanilla, and then we have these hybrid ones we talked about earlier, which I guess are like, I don't know, pistachio cream or something. There's like ordinary ones, and then there's these plaster ones.

35:19And part of the reason we're doing that, we're using multiple versions is, you know, some are easier to run on your laptop because it's like it's easier to get support for the thing. Some of them are easier to run in a big industrial data center. Some of them we release also because we want the research community to have options and be able to compare. Cool. And, you know, we have more coming. So we have, you know, the fact that we called currently the largest one small, perhaps portends that there is a large, there are larger ones on the way. And that is indeed the case. So we'll have a, we'll have larger ones, you know, where we're kind of, you know, polishing those right now, which, which, you know, they can do very, very sophisticated things.

36:03We're not going into the frontier model world, but we have pretty big ones. We have lots of great ones that run on your laptop very easily. The nano ones are interesting. So we have a 1 billion parameter and less models called Granite 4 Nano. Yeah. And those will run very nicely on a phone. So we've seen iPhones running at 90 tokens per second, which is faster than you can read.

36:31David Cox:like uh it's really i want a bunch of little ones on my phone on my iphone i've uh i've run a number over the ages and it's it's really interesting and i i keep wondering like i assume they'll continue to get smaller and and nano has been a big word this year and i'm like are we really going to get into like like sub nano pico pico femto we got two more got two more femto ato i think is one. Yeah, there's a number of stages as they get smaller. What kind of device do you use or what application do you use to run it on your phone, by the way? I know Corey uses, you use locally AI, right, Corey?

37:07David Cox:I use locally AI on my phone. Is that what you use, David? Yeah, I mean, there's different ones. I mean, I have a bunch of different ones. I'm an Android person. Oh, cool. And actually, it's cool because like running on a Snapdragon, second, there's a company, a startup called Next to AI. We don't know them. We don't pay them money. We don't have any relationship with them. They just went and said, hey, your model's cool. It's open. They took it. They optimized the crap out of it. Now it runs really well on lots of different devices. It runs in cars, runs in other places. And that's, again, that's the joy of the open ecosystem is you can just pick it up and let's all build stuff together.

37:45It's really cool. It's really cool.

37:47David Cox:What do you think is a common mistake you see when it comes to picking a model size for a task? I think people, especially people who are starting out creating agents, they will create something. And they'll often create this, like, long set of instructions that defines everything about what the thing's supposed to do. It gives a little backstory and, like, peps it up and says, hey, you're going to do great. some of these things actually by the way that there's people have proof that doesn't doesn't work like they're doing like pep talks and stuff but I saw somebody say like you will get a million dollars if you succeed in this task it's like what does that even mean?

38:26Or there's the one where they threaten you where they say like if you don't do this my entire family will die that's right yeah yeah Jesus anyway but they do that and then they it doesn't work for various reasons they're like I gotta get a bigger mall they go get a bigger mall or they go they juice it up that way and it's like actually just an ounce of you know care and how you craft the thing is worth 100 pounds of yeah being the model bigger because you end up being more expensive oftentimes the they'll also you'll you'll find that even just sending the prompt that you had through an llm to like clean it up and organize it will actually even improve things Yeah.

39:06But breaking things down into pieces also like that, that's super helpful. It also lets you be in control at various points if you're writing an agent. So I think that's the biggest thing. People, people confuse like when they should work on their agent and think about the structure with they should just go and get a bigger model. It's like bigger models are needed for this X, Y, or Z. I can't tell you how many things people have said, oh, you need bigger models for this. Yeah.

39:31David Cox:actually you just you know like not maybe go refine your 12 000 word prompt you've stuffed in there and exactly that's that's a big thing i see right now is that you know i was i was looking on reddit the other day at like uh one of the various subreddits that has things like that i noticed everything i clicked into had this prompt that was like 4 000 words and it looks nice and neat but when you really get to looking at it you realize there's tons of redundancies in here you're giving it more instructions than it's ever going to follow. You're clouding the context. Yeah, you're clouding the context.

40:05David Cox:And I think there's this temptation to confuse context with just more words. Yeah. Yeah. I mean, we knew how to maintain software. We learned that hard lessons over decades and decades, like most of a century. We don't know how to maintain a 10 ,000-word essay. No, it's weird. That's weird. And then people end up doing things like they'll capitalize words and they'll put exclamation points because it's like clearly something wasn't working. And then the problem is then you switch to another model because like whatever, a new model, a better model came out. So you want to jump to it. And then all that stuff you did is now just baggage.

40:41It doesn't help. It actually hurts. And you have to go and re-engineer it and do all this stuff. So that's actually the other thing that people do is they'll try it. They'll develop something with one model and they'll switch to another model and says, oh, that model isn't as good. It's like whatever you chose first is always the best one because you just optimized all that weird little like exclamation points and all that. So I think people are getting a sense of how these things work. It'll get better, right? People will get an intuition about this. They'll learn these lessons, in some cases the hard way.

41:10But people confuse consumer stuff with what you need to do with an agent behind the scenes, like how you interact with ChatGPT or Gemini. Wonderful thing. with that is itself an agent. Tons of software there. It's actually multiple models. So understanding all those things, people will get there, but those are the mistakes they make.

41:30David Cox:It is. And I think it leads to a lot of wasting compute that's being spent figuring out what you're asking it as opposed to answering your question sometimes and probably spending a lot of money as a result. So IBM is framing Granite as cost-efficient building blocks alongside larger reasoning models. I was wondering if you could tell us what that looks like. And is that a thing you think we're going to see more of in the future? I think you're going to see fewer and fewer cases of one big model doing everything. Yeah. Because that's just crazy. That's like, let's take all the things we do with motor vehicles in the city and let's have the biggest one do all the jobs.

42:15I have a tractor-trailer truck and it's just driving around. it's just like no you like scale it to your need like you have big trucks you have little trucks you have cars you got bicycles like like it's just crazy nobody would do that you the people think you're you're stupid or something so like like people prototype with the big thing and it's it would be too expensive to deploy and they kind of like chip off little pieces and say like hey actually i didn't need that i could just have this heterogeneous thing where i have other models that said also like the what's the definition of like it's just more and more things you can do with a smaller model.

42:48Yeah. Video just released a 30 billion parameter reasoning model, which is amazing. Yeah. You're going to keep seeing that happening. It's going to keep getting smaller and smaller and smaller. Smaller and smaller. And this inference scaling idea or the reasoning is really like, I spend more time to get more compute rather than like having a bigger model to get more compute. So you can trade this off dynamically. A small model can dynamically become big. I think it's just going to get to the point where, you know, it's going to saturate most things we're going to be able to do, we can do with smaller models.

43:15And then those tasks then that we can do with even smaller models, we'll do with even smaller models. We'll just keep optimizing it because why would you spend extra money? Why would you wait longer for an answer? It just doesn't make any sense. So that's the direction it's ultimately going to go.

43:29David Cox:Unless you just don't like money and time. You don't like money or the environment. You like burning electricity. Yeah. And you save budget, frankly, like your tech budget, your dollar budget, for those tasks where you really do maybe need a big model. Because laptops exist as well, you know, when you're dealing in giant data sets and all of those things. There may very well be the ability to go with a stronger model there by trimming elsewhere in your workflow. Absolutely. Absolutely. I use big models for, like, helping me with map papers and things. Like, those are very, like, high, you know, high skill sort of things.

44:10It's like, you know, esoteric, you know, kinds of knowledge. awesome. I'm not saying you should do that with a small box. You don't need your HR chatbot to know advanced physics. It's just not going to help. It's just not something you need. You're in trouble. You open it up to the possibility where people can jailbreak it and all of a sudden they're talking to the HR model about physics or worse things. Exactly. Do you have a thesis then on AGI? You treat these models as tools, do you believe that we will even ever achieve an AGI type system? Will it be multiple models tied together? What's your take?

44:52I don't believe there's any in principle barrier to us doing it. I don't think there's some like magic dust that needs to be there for it to be generally intelligent. I just mainly don't understand why we want it. Like why, why, why do we want that? It doesn't serve any purpose that I actually have. Like I can identify, Like, I just want things that help me do work that's, you know, things that are boring, things that are error prone. I want that taken care of in a way that I don't have to deal with it. I would like to keep doing thinking and stuff. And, you know, I would like humanity to be still centered.

45:25Yeah. Creative tasks. I would like to still keep having humans do those if possible. Like, I don't know what it's supposed to be there for. And there's almost like a weird kind of cult-like vibe sometimes you get from folks. like they're ushering in the the awakening of you know the that which comes next or something and it it just it just it just seems a little distasteful it seems kind of counter to this human notion of like hey like we have tools we have like we've been having tools since we like picked up a rock and started banging something like now we have computers you know like now we have ai like it's it's like that's the extension of that and that seems better i don't necessarily want there to be, you know, the new God or something.

46:08But I also don't think there's anything in principle that would stop us from doing, you know, very intelligent machines. And people sometimes also say like, oh, if only imagine you could have 200 PhDs working for you all the time. And I'm like, I have 200 PhDs working for me all the time. That causes more problems than it solves. Like, it's...

46:28David Cox:I want someone under me to manage all those PhDs. That's what I'm talking about. Yeah, seriously, figure out the management structure, let me know. But like they think that that's going to solve some problem and it just kind of like moves the problems around in many cases. And those weren't really the problems to begin with. I do think there are going to be fields of science where like, hey, like having an army of brilliant physicist bots, if we can get there. Yeah. That could be very exciting. But for most things, we just don't need it. One of the things that Corey and I are really excited about is like voice as a new interface.

47:02So I know you've mentioned that maybe the chat interface isn't good for every tool use case. Like perhaps you just want it to work like functions that you call. But what about a case where you do want to chat naturally with your computer and not have to even do typing as an interface? Or is Granite working on something? Absolutely. And to be 100 % clear, when we're using it as an interface and we know we're chatting, like this is great. Like, of course, like that's a very good way. And you do want to be able to tune the persona and things like that. And actually, you know, one of the things that we've been playing with is having, you know, all of our models are small and fast.

47:39We've also been playing with modular speech to speech, where you can just like the most natural way to have these conversations. like this isn't us mistreating a computer as a thing we talk to. We wanted to have a conversational interface. You just talk to it and it talks back. And it can be really fun and compelling. One of the reasons why businesses don't deploy these things like all that often, the speech-to-speech models, there's lots of speech-to-speech models. But the problem is if you don't have an opportunity to check what it's going to say before it says it, then, you know, like normally you'd put a guardrail on it and you check what it says and then before it says it.

48:17But if you have these models where they've actually been trained to just take in speech and just put out speech and you have no say in it, that's a little bit harder. We've created modular ones where basically we can be in parallel doing guardrail. Like we can be saying like, hey, is this thing going to say something biased or against company policy? We can be checking and checking for hallucinations. And then before it even begins to speak, we can stop it. We can say, hey, oh, it was going to do something wrong and we can spend more time to generate. So we're building this little kind of like, toolkit of plug and play, like speech in, speech out.

48:50And we were releasing it recently at NeurIPS, which is the big AI conference that happens every year. This year it happened in San Diego. And we had a demo at our booth where Luis Ostras, who's a brilliant guy, director in our lab, he leads many things, including the speech effort, did a demo live on the floor, basically having conversation with Granite about how Granite works. So we have a video actually, if you want to see, like, understand a little bit about how Granite was put together, like from the authority, which is Granite itself, you know, it has a little conversation, which is a lot of fun.

49:29And then at the end he writes, you know, he writes a little poem based on the conversation that he had with the folks in the audience. So it's just fun when you can get some of these things to work and And they're kind of fluid like that. They're all running on one GPU, very small, very light. Wow. Very trustworthy. So, David, if we look back on this conversation a year from now, three years from now, what do you hope we can say IBM got right? Yeah, I hope we get to the place where people can deploy these things and make everything better, like make businesses run more efficiently, make customer service easier, make all those things that we would rather ignore.

50:13I hope we can do that responsibly and safely and we can just forget about it. So maybe a succinct way of saying that is, you know, when stuff, technology is new, it's exciting. Like when electricity first came out, it was very exciting. What does that mean? Like you had Tesla on one hand with alternating current, yet Edison on the other with direct current. Direct current was the wrong answer. It burned houses down. They caught on fire spontaneously. It was very exciting. Not in a good way. Now electricity is like, it's boring. It's boring. It's like it just melted into the background. Things, you know, technologies go through those waves where they're exciting, but then they just melt into the background.

50:49And as much as I love having attention and like all this stuff happening in the field that I work in and I've worked in my entire life, my adult life, I would love if IBM made it boring. Like it's just like it's uncontroversial. You can deploy it. There's no risks. There's no like weird stories about people getting their hard drives reduced that we put in the structure. Blackmailed. Yeah. Or like, you know, I had like a chat bot giving the wrong policy, which has to be honored because the judge rules that it has to be honored. That's real. Yeah. This has all happened. And I would love it if it became so controllable and controlled and indeed boring.

51:25That would be great. And then we have excitement somewhere else. We can save that for some other place. David, this was fantastic. Thanks for making hybrid architectures and enterprise constraints actually interesting. That is no small feat. Very good. To your point, you want it to be boring, but you actually make it really fun to talk about. So I appreciate that.

51:49David Cox:You do. You do. It's fascinating. So if you're listening and you want to go try Granite 4.0 for yourself, we'll link to the Hugging Face archives where they've got some spaces set up. And is there anything else we should share out, David? Nope. You can check it out. You can run it right in your web browser locally on your computer. Nothing installed. Just go to a browser. It'll run in your browser. Amazing. That's awesome. Awesome. Is that true for the speech model as well? or how do you run that on a computer? That one, no. But we don't have a space for that one yet. But you can do like, you can document model we have out there.

52:24We're putting out more and more of those. Yeah, cool.

52:27David Cox:Well, if you haven't yet, please take a moment to like and subscribe. Really helps us continue bringing you guests like this who are innovating at the frontier edge of AI today. Also, make sure you stop by at theneuron.ai to sign up for our daily newsletter and join 600 ,000 or so others who read it every morning. if you're building a

52:50and on that note

52:51David Cox:we'll see you next time farewell for now humans

From the publisher

IBM just released Granite 4.0, a new family of open language models designed to be fast, memory-efficient, and enterprise-ready — and it represents a very different philosophy from today’s frontier AI race.


In this episode of The Neuron, IBM Research’s David Cox joins us to unpack why IBM treats AI models as tools rather than entities, how hybrid architectures dramatically reduce memory and cost, and why openness, transparency, and external audits matter more than ever for real-world deployment.


We dive into long-context efficiency, agent safety, LoRA adapters, on-device AI, voice interfaces, and why the future of AI may look a lot more boring — in the best possible way.


If you’re building AI systems for production, agents, or enterprise workflows, this conversation is required listening.


Subscribe to The Neuron newsletter for more interviews with the leaders shaping the future of work and AI: https://theneuron.ai

More from The Neuron: AI Explained

All 106 episodes
Why IBM Wants AI to Be BoringThe Neuron: AI Explained · 53 min
Listen in VO