In short
IBM’s Granite small AI models (2B/8B) and why IBM is “betting everything” on directly training small models (Granite 3.2/3.3) rather than distilling from larger models, plus how RL, inference-time scaling, and modular runtimes (LoRA/activated LoRA) can match much larger models’ code/math performance while improving cost, safety alignment, and deployability across hardware.
Guest
Sriram Raghavan, IBM (Granite/Watson X) researcher/leader focused on small language models, training pipelines, data quality, and inference/runtime techniques; discusses Granite 3.x releases and open models.
Key claims
- Distillation improves narrow benchmarks but can reduce broader capabilities and safety/alignment.
- Direct RL training on 2B/8B avoids those tradeoffs and preserves base performance.
- Granite 3.3 with RL + inference scaling can match GPT-4–class code/math performance despite being 100x smaller.
- Model “currency” is shifting from parameter count to memory/latency; discusses mixture-of-experts and state-space (Mamba-like) hybrids.
- Modular “generative computing” runtimes (LoRA plugins, checks, routing) reduce the need for huge agent prompts.
Notable examples
- Granite 3.3 matching GPT-4o/“o3.5”-level code/math benchmarks.
- Data Prep Kit (DPK) and IBM’s NiceWeb dataset with published cleaning recipes.
- Particle filtering as an inference-scaling technique.
- Activated LoRA reuses KV cache for 2–3x gains.
- Bamba 2 (state-space hybrid) released with CMU/Princeton/IBM Research.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOGranite 3.2 and 3.3 Models Explained
0:46 to 2:49
Discussion on the unique features and advantages of the 2B and 8B parameter models.
“So Craig, first of all, good to see you again.”
The Shift from Distillation to Direct Training
2:50 to 4:50
Examination of the rationale behind training smaller models directly without distillation.
“Because there's more you can do, you open up AI to run more cheaply in wider environments.”
Reflections on Model Training Costs
6:33 to 8:19
Insights into the evolution of model training and the impracticality of large models.
“is a registered broker-dealer and member of FINLA, NFA, and SIPC.”
The Importance of Data Quality and Model Architecture
8:20 to 10:52
Discussion on the improvements in data quality and architectural innovations in AI.
“But it played out because of a combination of factors.”
Collaborative Efforts in Data Preparation
10:53 to 13:32
Exploration of community-driven initiatives for data cleansing and sharing best practices.
“from pure pre-training to all of these other stages that have gone.”
Future Directions with Mixture of Experts
13:33 to 14:00
Discussion on potential advancements in AI model architecture and the concept of mixture of experts.
“And also remember, because the training stages have become more, there is now data sets at every stage.”
Architectural Innovations in AI Models
14:00 to 16:44
Explore the advancements in AI model architectures like mixture of experts and state space models.
“But is there an advantage once you've worked out this training strategy or pipeline that then you can start building bigger models?”
Balancing Training and Overfitting in AI
16:44 to 19:55
Learn about strategies to avoid overfitting in AI models while balancing various tasks.
“away with H 100 can I work on an A 100 can I run on an L 40s all of them start to become important questions.”
Exploration and Backtracking in AI
19:55 to 22:28
Understand the importance of exploration and backtracking in teaching AI models complex tasks.
“In certain other cases I will need an external program like a trained reward model to teach the model how to explore.”
Inference Scaling and Its Implications
22:28 to 28:00
Discover how inference scaling can optimize model performance and the trade-offs between small and large models.
“Because we had trained the model in our RL training direct on top of the model with enough high quality long chain of thought data, then it had a very good understanding of code and math.”
Show all 23 chapters
The Shift to Smaller AI Models
28:00 to 29:30
Learn about the advantages of smaller AI models over larger ones.
“you need a big model to have seen a large amount of data to be able to do that.”
Generative Computing: A New Paradigm
29:30 to 31:00
Explore how generative computing is transforming AI applications.
“You will see now order of magnitude reduction in model performance.”
Introduction of LoRa Adapters
31:00 to 32:50
Understand the role of LoRa adapters in enhancing model functionality.
“you use rag to give the model enterprise that's not appearing I think we are at a point where we are now, all of those are true.”
Runtime and Model Interaction
32:50 to 34:20
Discuss how runtimes enhance interaction with AI models.
“along actually with granite 3.3 instead of what we call LoRa adapters right these do one thing very very well so they you know there was a LoRa adapter that does hallucination detection.”
Programming Requirements for AI Applications
34:20 to 36:10
Discover how developers can specify requirements for AI outputs efficiently.
“So you no longer are talking raw to the model.”
Innovations in AI Efficiency
36:10 to 37:50
Learn about new technologies improving AI model efficiency.
“You can iterate through, and you will achieve the same effect.”
Small Models and IBM's Strategy
37:50 to 39:20
Explore IBM's strategic focus on small AI models for enterprise adoption.
“LoRa's, activated LoRa's, I'm sure there are other innovations waiting that will allow us to write really complicated things.”
The Future of AI Model Development
39:20 to 42:00
Discuss the implications of small models on future AI development cycles.
“Is that part of the strategy with these small models?”
Continuous Learning and Modularity in AI Models
42:00 to 45:50
Learn about the concept of continuous learning in AI and how modular components like LoRa can help adapt models in real-time.
“whose release cycles tend to be six, nine months a year.”
Uncertainty Quantification in AI
45:50 to 48:58
Discover the importance of uncertainty quantification for AI systems to improve their self-awareness and adaptability.
“And their small is easier to update, that's about it.”
Innovations in Online Learning for AI
48:58 to 52:49
Explore the challenges and innovations in achieving online learning in AI through modular methods and LoRa integration.
“if inference time weight adjustment is possible, would it happen at layer-specific updates?”
IBM's Focus on Data Management and AI Agents
52:49 to 56:00
Understand IBM's strategic focus on data management and the development of specialized AI agents for enterprise applications.
“Is the Granite series, is that the main thrust on generative AI right now?”
Building Specialized AI Agents at IBM
56:00 to 58:12
Learn about IBM's focus on developing AI agents for specific business domains.
“So imagine Craig you're an IBM employee.”
Transcript
Automatic transcript. May contain errors.0:00Wonderful to talk to you again. We've spoken on the phone. I met you here at Yorktown Heights a few months ago. At that time, you were getting ready to release Granite 3.2, and I wanted to follow up on that, but I was interested in what you guys were doing. In particular, and I think it's still the case, you have a 2 billion parameter model and an 8 billion parameter model in the 3.2 series. And what was unique about those models, from what I understood, is that you train them directly. They're not distilled from a larger model. So can you talk a little bit about why that's an advantage? Yep. Yep, absolutely.
0:47So Craig, first of all, good to see you again. And welcome to Yorktown. So, we'll maybe wind back a little bit to our conversation. I think that was around the time when it's unseen the DeepSeek models released. And distilling, well, distilling is an old technique. It had come back to the fore because DeepSeek, in addition to releasing their big model, had also shown this idea that you could take smaller models. I think they had experimented with Lama and Mistral, smaller models. and said that using the big model, I could distill those Lama and Mistral models down and get their performance on specific benchmarks like code and math benchmark to really shine well beyond the original capabilities of those models.
1:29So that's why distilling was in four. So that's when we had the discussion and that was with 3.2 and we have continued that actually and released 3.3 now in the same approach. So there were two reasons why we wanted to explore an alternate approach. So one is, while distilling allows you to optimize, especially if you, because when you distill, you are basically essentially having this bigger model teach that smaller model to do better at something. That tends to do, make that smaller model do better at that thing, but it does break some of the base training and alignment of that model. So what we had shown was, yes, absolutely, LAMA and Mistral models do better at the code and match benchmarks than they were originally doing on their own.
2:17I believe they showed it with the LAMA 8B, I think similar to Mistral models. But you do then lose the broader capabilities of the model. So their bench performance on broader benchmarks goes down. And more importantly, it also often, and we validate with measurements, takes away some of the safety alignment of the models. So what we were doing was to say, we certainly see the logic of that. But as you know, our overall strategy with Grandad has been really focused around small language models. And what can you do with smaller models and how far can you push them? Because there's more you can do, you open up AI to run more cheaply in wider environments.
2:57And as you know, we've been doubling down on 2 billion and 8 billion dense models. And we started that with our 3.0 series. So we said, in this particular case, can we do RL, reinforcement learning, on our 2 and 8 billion models directly? And furthermore, on top of that, how much can we push with inference scaling or runtime scaling? Now the reason to do that is we can do that without, by avoiding the two pitfalls that I said before. I can do it by saying can I do this in a way that yes it gets better at what I'm teaching it, but it still holds on to the overall base performance. And we were able, we showed that you could do that with 3.2 and we continued that journey with 3.3.
3:42We further, and I don't know whether we talked about it at that time, we had also been innovating on inference scaling techniques on top. So as a very concrete example, our Granite 3.3 model with a combination of the RL on top plus inference scaling using some new techniques, and we can talk about those techniques, can match the performance in code and MAC benchmarks of GPT-4 row CLOD 3.5 models that are two orders of magnitude bigger in size. So that's sort of the journey that we are on. And the reason we like that is A, it addresses those two disadvantages. Second, we control the origins of the data that we're doing.
4:22When you're distilling from some other model, you're always beholden to the origins of the data. And as you know, we are very focused at IBM on making sure our granite models have lineage data that we control, we have curated and we have gone through the vetting process. So this path gives us an organic way using data that we can fully speak to, are aware of, have curated to produce models with very, very, very high quality performance. This episode is brought to you by Tasty Trade. On Eye on AI, we talk a lot about how artificial intelligence is changing how people analyze information, spot patterns and make more informed decisions.
5:05Markets are no different. The edge increasingly comes from having the right tools, the right data, and the ability to understand risk clearly. That's one of the reasons I like what Tasty Trade is building. With Tasty Trade, you can trade stocks, options, futures, and crypto all in one platform with low commissions, including zero commissions on stocks and crypto. so you keep more of what you earn. The platform is packed with advanced charting tools, backtesting, strategy selection, and risk analysis tools that help you think in probabilities rather than guesses. They've also introduced an AI-powered search feature that can help you discover symbols aligned with your interests, which is a smart way to explore markets more intentionally.
6:02For active traders, there are tools like Active Trader Mode, One-Click Trading, and Smart Order Tracking. And if you're still learning, Tasty Trade offers dozens of free educational courses, plus live support from their trade desk reps during trading hours. If you're serious about trading in a world increasingly shaped by technology, check out Tasty Trade. Visit tastytrade.com to start your trading journey today. I'm going to myself. Tasty Trade Inc. is a registered broker-dealer and member of FINLA, NFA, and SIPC. Yeah, and that idea of direct training smaller models, just in the evolution of these models, Is it the case that if you knew what you know now, back before OpenAI trained GPT-2 and 3 and 4, that you wouldn't need to spend that much money on training that large a model?
7:17I mean, not that there aren't some advantages to having a large model, but it just took building those really big models to realize that you can do it with less. Sometimes it's the arc of technology. Yeah. Or, you know, you can call it hindsight 2020. But I do think that it was, the field needed to go through at some level. Now, it was intuitively clear to a number of us. And we had been talking about efficiency, cost and size from the get-go. And I'm talking 2023 when we released Watson X. Even at that time, we were talking about that being going to be important. it was intuitively it was clear that if you know what you're going to do and it doesn't have to be one thing it could be a collection of tasks that an enterprise needs versus I'm going to give you a model that's going to do many many many things including things I have no idea what you're going to ask it to do then certainly possible to do it with much much smaller model but I do think that that intuition played out and in fact it played out maybe I would say even faster than you could have imagined.
8:20But it played out because of a combination of factors. So one is, A, we certainly got better as a field, as a community around data quality. So that allowed us to say, you know what? I can now do today, I mean today in our training process, Granite 3.3, halfway into its training, was way better than Granite 3.0. Halfway into its training. Why? Because we have done so much data quality work that now when you are digesting tokens to put into the model, you get better. So data quality has certainly improved as a community. Second, there have certainly been architecture innovations. So now we know how to more efficiently package what you can in smaller number of parameters.
9:03And I think the third to your point is, I think we have gotten better at understanding what are the real requirements for enterprises. And since our focus is on enterprise AI, I think these three combination of factors has led to where we are today. And as you rightly point out, there are always scenarios where larger models are needed. We know, for example, for extended reasoning tasks, we know larger models do better. But that's where this inference scaling idea we like. Because the advantage with inference scaling is it lets you take a small model, spin compute to make it behave like a larger model.
9:39And I can turn that on demand. But if I've built a larger model, it's not easy to shrink it. I'm always paying that cost. So we like this idea of continue to push the smaller model, make it better and better and better. Then always you have the luxury to spend more compute time to make it do even better when it really, really matters and it's an important question. And that size, the two billion or eight billion, that is a direct correlate to the size of the training data. Is that right? No, actually the training data has remained quite... So, in fact, I would say the other way, which is we had bigger models that were sometimes consuming smaller data early on.
10:21We have managed to even pack more data into small models. Now, what has happened is, in fact, I think the data has gone to occur. So early on, it was about... The whole conversation was about two things. size of the model and the amount of pre-training data. Now it's become architecture of the model and the quality of the data. So that's sort of the shift that has happened. Now you do still put a fair amount. I mean, these models still take 6, 8, 10 trillion tokens worth of data, but they are certainly much, much more quality aware. I think the second is the focus has also shifted a lot from pure pre-training to all of these other stages that have gone.
11:00So if you sort of look at what has happened with training, We had this very simplistic view of you do pre-training, which is the self-supervised learning, and then you do post-training. And post-training typically used to be basically supervised fine-tuning or instruction tuning. Now, if you look at the pipelines of what us and others are doing, there's obviously pre-training. But then there's like this mid-training where you do some things like prepare the model, make it ready to receive RL. That involves some amount of supervised fine-tuning. Then you do specific actions to make the model good at chain of thought.
11:35Then you do the RL phase. Then you do supervised fine tuning. You might do a couple of rounds of RL. So the stages have become more sophisticated. And almost to the point where you could argue that base pre-training, which used to be the biggest focus early on, is almost, I'm going to use the word commoditized. I'm going to use that in this context. Yeah, everybody does pre-training. Yeah. And to improve the quality of the training data, how difficult the process is that? It is. It is probably the one where a lot of work has happened again in the open community. We've contributed. Others have contributed.
12:15In fact, we've created open source projects, a project called Data Prep Kit, DPK. It's born out of something we use to allow the community to create and share recipes for cleaning pipelines. Because there isn't necessarily one way to do it. I may make a certain set of decisions in my model training pipeline because I know what I'm doing. Content that I'm actually discarding may actually be important for you. Imagine you're building a model explicitly to discard poisonous content or hateful content. I'm going to throw it away because I don't want my model to see. You might actually want it because that's what you want to build your detectives, right?
12:52So we know that data cleansing requirements, processing requirements are different. So we're starting to share the infrastructure and the pipelines to do that. And then websites, you know, communities are releasing data sets. So I think that has been an important focus. In fact, interestingly, we even released from IBM, and we can send you the information, a data set we call NiceWeb. And what we tried to do is to show, hey, here is the cleaning recipe we used. Have at it. You can improve the cleaning recipe, take it, adapt it. So that's now become a currency of sharing in the open environment.
13:27So we went from just sharing data sets to sharing recipes that now you can adapt the recipe. So I think that's a fascinating journey. And also remember, because the training stages have become more, there is now data sets at every stage. A lot of the focus used to be on just pre-training data. Yeah. Now instruction tuning data, RL data, all of that is starting to become important. And it's a very, very vibrant community. That's what's so fascinating about it. Yeah. And do you intend then to scale up the 8 billion, even though you don't want to go presumably to 400 billion or something? But is there an advantage once you've worked out this training strategy or pipeline that then you can start building bigger models?
14:14I mean, is there any advantage to that? We can, but I think one of the things we're making is we're also now doubling down on mixture of experts architectures. And actually, let me talk about sort of two architecture innovations. I almost think that, I think we're at a point where the number of parameters as a model starts to not become the right descriptive element. Because what happens now with these architecture innovations is that simple number that we use to track sizes probably is now underserves the purpose of figuring out what model choices are. So one is with mixture of experts now you have two numbers.
14:52You have the overall parameters used and then the actual size of the individual active parameters. We're also now and we actually worked with Princeton and CMU on bringing states-based models. Now, state space models are models that try to address the quadratic computation bottleneck of attention, right, with state space approaches. So we actually worked with the researchers across CMU and Princeton, including some of the researchers who had originally invented space space model to look at hybrid architectures, where you can combine classical attention or not classical, the traditional attention and state space together.
15:31So you get the efficiency benefits of state space while preserving the overall accuracy quality, right? We release that in the open. So now what's happening is I think model parameters are going to not become the most important currency. A better metric might be what is the memory requirement of the model? And you characterize that by asking, okay, on average, let's say I have this many user sessions with this context length. I want to run against the model. How much memory do you need? I think that is a much more useful functional characteristic because you might have, I don't know, 50 billion parameters.
16:11I might have 38 billion parameters. But the choice you made with how much state space you are using, what mixture architecture, number of exports you used, your 50 billion with your choices might actually take lesser memory than my 34 billion with a different set of choices. So I think very soon we will recognize that yes there is a simplicity to using number of parameters to talk it's just easy to understand but what's going to matter when you start making model choices is really going to be memory requirement is more and more that's what's determining your latency performance do I need a H 200 can I get away with H 100 can I work on an A 100 can I run on an L 40s all of them start to become important questions.
16:55And on state-based, is that Mamba that you're talking about? So we released a model called actually Bamba 2. So it's in that family. Right. Inspired by Mamba, and we'll send you the information around it. But Bamba 2 was jointly done across CMU, Princeton, IBM Research, and we released that in the open. We are bringing some of that learning into the next generation of Granite 4 that we are working on, which will be both mixture of experts and have some of that architecture. So lots of innovation. Yeah. And the advantage is that, look, the currency, as I said, is memory. And so the question you're asking yourself is, what's the size of the KV cache when your model is running?
17:32That's like, you know, the most important performance question. Yeah. In enhancing the reasoning through direct training, and I'm reading a question now. how did you avoid overfitting to logic heavy tasks like math and coding at the expense of general? Yeah, so it's always the usual trick in AI which is... And actually in that, can you talk about thought preference optimization? Right, right, sure. So I think the game is always balancing the task you want with constantly reminding the model about the other things you have to do. So there's a training procedure that constantly credits the model both for what you want it to do, but keeps reminding the model for all of the other things you want.
18:25There's a use of a replay buffer from past data. That's where some of the balancing act comes. Also remember that this idea of training for code and Mac, the reason we all go there first, is those are tasks for which you can write verifiers easily. Right. It continues to remain an interesting research question, which is, how are we going to scale this to other domains, where verifiers will take work to write, possible to write, we can certainly do that. Are there scenarios in which you deploy models and then learn those chains of thought by actually observing humans in action? Like imagine I'm looking to debug IT systems, which is an area I'm interested in.
19:12We're interested in as other company. Now the chain of thought, the sequence of steps that an SRE uses to look at a problem. Let me go look here. Oh, that isn't it. Let me go look there. Well, that isn't it. If this and this are not true, let me go look there. That's a chain of thought. That's backtracking. So, teaching that to a model is a different task. Now, the model will understand how to do this broad idea of exploration and backtracking using code and math. But how to backtrack, how to explore is a very domain specific thing. And that is something we are going to have to continue to teach the model.
19:49So, back to your question, this balancing act is usually the usual trick in AI, which is create your objective function your data so that you are not overfitting for the task you're constantly balancing and then I think this whole space of RL I would characterize it as there are sort of two variables you're going to see a lot of innovation on one variable is how am I getting the model to do the exploration am I doing it by generating a lot of synthetic data in In some domains that's the best way, which means I know I can synthetically generate chains of thought and I do training of the model data driven.
20:28In certain other cases I will need an external program like a trained reward model to teach the model how to explore. So that's going to be one vector of innovation. The second vector of innovation is this whole verifier. So how you do exploration and how you do verification I think are going to be the game and you're going to see thought preference optimization, a process reward model. All of these things, I think, are going to evolve in very interesting ways. One example of work actually our innovation team in Red Hat jointly with IBM Research is this exploration called particle filtering. Called, I'm sorry.
21:07Called particle filtering. It was an inference scaling technique. And the intuition for the idea is look can I think of the task of that I'm trying to teach the model as a probabilistic inference borrow ideas from probabilistic inference and what you end up doing is you keep a lot of parts active so instead of always picking the best at any given time you keep a lot of and that's why it's called particles because it's like you're exploring at any given time there is a frontier of particles, you keep a number of them active at any given point of time. Periodically filter some, explore the remaining particles.
21:48And by formulating it, by taking ideas from, you know, well understood, probably in inference, you get a very efficient scaling technique, right? Inference scaling technique. So I just think that this is going to be one of those rich spaces. What we have seen is, for things like code and math, the data side, right? Because you're able to now there's lots of data sets out there able to generate data even simple majority voting pushes a long way. Now it works very well when you want to hit benchmarks. It's not clear that majority voting will necessarily be sufficient for every other domain but in domains like code and math the result that I showed you where our 3.3 model went up to PORO it were able to do that with just majority voting.
22:29Why? Because we had trained the model in our RL training direct on top of the model with enough high quality long chain of thought data, then it had a very good understanding of code and math. Last comment I will make is, I did say that, yeah, pre-training is getting commoditized, right? And all the interesting innovation is downstream. But the amount of tokens of certain kinds you put in pre-training still matters. What do I mean by that? Is if I want to build using all of these kinds of innovations a high quality code or math model I better have put a non-trivial amount of code and match data to pre-train if I don't do that then this ain't going to work so there are some dependencies but I think this is getting to a point where a lot of the data mixtures for pre-training are sort of getting standardized in the community so I think everybody is adopting you know roughly the same.
23:27Can you talk about inference time scaling? In particular, as I recall, I'm not sure in 3.3, but you allow developers to turn it on and off. Correct. In that, you know, as a consumer, when you use these models on your desktop, you can choose whether you want reasoning or not. So what's the difference? I mean it's this is through an API it's switching it on and off in the code but... Right so the backend implementation right so is the differences between do I have two different models and when you select I'm actually routing you to two different models or my same model can turn it on or off on the model.
24:17That's the real difference and it makes a difference because running two versus one is cost And I have to drift. So that's the difference. But back to inference scaling. So I give you one example of this particle digging approach. There's lots of other approaches getting implemented. I think that inference scaling, depending on... I like to think of it as the... My simple mapping is how much compute are you spending at runtime. Now you could spend compute by taking one model asking it to run that many different times and do something. You can take one model do this a kind of exploration. I also put inference scaling under the category of I want to answer a user question.
25:03I'm going to take multiple models combine their result. That also in some respects you are scaling because to answer that one user query you are now spending more compute. So I think of inference scaling as this broader category that says if your model is primed well, then you can get stuff out of the model that are much better than simply asking the model to give its best answer. When the model gives the best answer, it's making a set of decisions throughout the network. And you're basically saying, okay, that's your best guess. Now try harder, evaluate more possibilities. I think that is the general idea of inference scaling.
25:43And the reason I think there's going to be a fertile area of research is, A, small models have an advantage here because it's possible to make a small model behave like a big model by inference scaling. And I pay the cost only when I need to. Whereas it isn't possible to go the other way around. So everything being equal, if I told you, you have a complicated task to do. You want to do that complicated task once every, you know, one out of ten tasks you want to do. Nine out of the ten tasks is not a complicated task. And if I can make sure that I can get you the expected accuracy of that complicated task, and I can do that with a 8 billion model or a 10 billion model with inference scaling, I can optimize this to the point where its cost is as good or cheaper, way cheaper in practice, than a big model that you end up having to host for all of the nine tasks where you don't need the model.
26:43So that's why I believe that this cost and optimization is going to drive this move towards small model asked many times better than a big model asked once. The big model probably has the capacity to give the correct answer right off the bat. But it's not important if I can figure it out because you're all hidden behind an API, I get the answer back. The second point is the small versus big and this is a different dimension not just an inference scaling. One of the reasons sometimes you need a bigger model is especially when you build more complicated applications like agents. Your prompts become really really big.
27:21Now what happens is if you look under those prompts, when you take a big complicated agent prompt you you peek under the covers and sometimes these things run into pages and pages like an English essay. You will see on the prompt instructions to the model, requirements, security, thou shall not do this, thou shall do this. It will have formatting. It will insist that this must come before this, this must come after this. So it's really a program. So it's a program written in English. And now you can see why you really, if you want the model to pass through all of that, you need a big model to have seen a large amount of data to be able to do that.
28:06But if you take that simple leap of saying, look, this one is not something an end user is going to do. This is a developer task. End users working with a chat GP data interface want to use natural language, I get it. This is a developer task. So we see a shift from massive prompt engineering behaving as if it's a software engineering task to let me actually turn those into actual programming tasks. When you start to break it down, hey I'm going to declare some instructions. I'm going to declare the requirements for the model. Here are do's and don'ts for the model. You start to decompose it, you can get away with a much, much smaller language model.
28:46So with a little bit of procedural logic on top of a small language model, you can get as good performance as a large model. So I see two vectors pushing towards the fact that small languages model can get you what you want in many, many cases. One is this vector of inference scaling, which is a built in. The developer doesn't see it. The model does it. The other is where the developer is involved in saying, you know what? I know what I want you to do. I'm just going to break it down. And in fact, this is where you can even do checks outside the model. You can say, hey, as a developer, I know I wanted the last piece of the document to always have my name in the email signature.
Read the full transcript
29:27It's a very simple check to do. If it doesn't show up, go back to the model, try again. These things can be done. You will see now order of magnitude reduction in model performance. So we certainly see these two vectors continuing to push the needle on how small models are going to do. Today, because of the way the field has evolved, prompts have become sort of this only lingua franca. But we're at a point where we're building more and more sophisticated applications. If you're only building a Q &A application, prompt is okay. If you're building a full-blown agent, talking to tools, making planning calls, iterating, reifying, I think you're going to see programming abstractions.
30:06And we're very actively working on some of those programming abstractions. And that's done by the developer in the code of the application. At the application. But we are going to give them now programming frameworks. So we have started to use the word, by the way, generative computing. So we keep talking generative AI. We really believe we are on this cusp, and you will hear more about this, I think, when you're there. You're going to see us talk about generative computing, because I think we are at the shift where we went from, hey this cool self supervised model LLM cool what can it do then we went from that to okay you know what a model is part of an entire platform so that's where you started having platforms from us and others we built Watson X where models are important but you need all this plumbing around the models we talked a lot about models as a data representation we said it's important to get enterprise content into those models that's why techniques like InstraClab you use rag to give the model enterprise that's not appearing I think we are at a point where we are now, all of those are true.
31:09It's a powerful self-supervised representation. It's core to an AI platform. It is a data representation. But I think we are at the point of saying a model is now a new type of computing element. If it's a new type of computing element, a computing element has abstractions. It has a runtime. It has ways to speak to it. And I think that is what we see starting to emerge. I think people are very quickly realizing. So we think that just as we talk about classical computing and quantum computing, we are very comfortable now starting to use the word generative computing. Now only difference is that unlike quantum, it's not an entirely new hardware off on the side.
31:50It's a hardware we understand. But I think treating the model as a computing element now allows you to ask, how do I program this computing element? A programming a compute element doesn't have to mean writing five pages of English. Programming a computing element can mean here are the abstractions. So that's where we are pulling and it's motivated by what we just discussed. If I can do that I can do very very well with much much smaller models and that drives efficiency. Yeah and when you say generative computing is there a model that's that's generating the program I mean actual code to do these things?
32:30No so we are using generated computing to what we described which is a program that you use to work with the model but not as if you just have to give it a prompt but you can actually use instruction so as a concrete example let me let me give you a flavor of water generative computing program could look like so we released along actually with granite 3.3 instead of what we call LoRa adapters right these do one thing very very well so they you know there was a LoRa adapter that does hallucination detection. There's a LoRa adapter that can rewrite queries better for better search performance. There is an adapter that does uncertainty quantification, which says these are all trained adapters for those specific things, right?
33:14So you can think of these LoRa adapters almost like plugins stuck on top of Granite 3.3. So you take Granite 3.3, you can have three, four, five adapters enabled. Now a generative computing program will actually, where a program you say what I'm building a rag application and I want to turn on or off one or more of these LoRa adapters. So that's what we mean by generator point. Now why is this interesting? This is interesting because I could do this whole thing without these individual LoRa adapters with a big model and in English say I want you to do hallucination check. But now that big model has to be trained to get all of this done.
33:52It has to rewrite the query, it has to do hallucination detection, it has to do uncertainty quantification. But now by decomposing it into a good base model and individual LoRa adapters, the programmer controls. They stitch a workflow together, turn it on and off when needed. I'm working in an environment where uncertainty score is not important, I don't need that LoRa adapter. So that's what we mean by a generative computing program. So the model is really wrapped by a runtime. So you no longer are talking raw to the model. You're talking to a runtime that wraps the model. The runtime exposes abstraction that lets you on and off functionality.
34:31And then the runtime enables or disables these lower adapters. Eventually you're running an inference through. Finally that's the operation. But you are setting up that inference based on these flags and other controls. Right and turning on and off these elements, turning on and off inference time compute for example. Is that automatic within the model does it decide? The runtime. Now the runtime associated with the model does it. Right. It's not the developer in advance. No. Yeah. No. So you just set up that environment and the developer comes in and uses this abstractions. So for example take this one.
35:13You have a requirement to you have a requirement that your output looks a certain way. The developer can say here is a little function that's going to check that you meet the requirement right. It can be a little piece of a program. So think about it if you didn't do it this way then you have to teach that model how to do this requirement. That's why you end up writing these laborious English using examples that show this is the right this is wrong because that's the way for the model to learn. On the other hand all of that is useful if the model is directly exposed to an end consumer in a consumer use case.
35:47We're talking about enterprise applications that developers are writing. So if there is a simple check that says, this document's output must always end with date and IBM confidential at the end. That's a simple check that any programmer in high school can write. You declare a little function. Okay, you write your generative computing program. You can iterate through, and you will achieve the same effect. So that's what we mean by the generative computing program. What's interesting is there are now more and more efficient ways to do this. So we talked about LoRa adapters. One of the challenges with LoRa adapters is, yes, they allow you to build these modular pieces of functionality.
36:29But anytime you invoke a LoRa adapter, it's like you're processing a set of tokens. Then you invoke a LoRa adapter. While processing the set of tokens, you would have generated a KV cache for the base model. Soon as you invoke the LoRa adapter, you'll have to throw out that KV cache because you'll have to redo the KV cache because the LoRa adapter has been trained on its own. So you'll have to repopulate that. So you've been developing an interesting technology called activated LoRa. And the goal of the activated LoRa is, can I share KV cache with the base model? So you can turn me on and I can actually reuse the work that was done by the base model in processing some amount of tokens.
37:10then just do my thing that I need to do, move on and a base model can continue to move on. So this is now squeezing out even more. We have seen activated LoRa's can give 2-3x advantages between dedicated LoRa's. So you see where I am going with all of these things. They are all about finding powerful ways to bring down. So we are very, how to say, very convinced that abstractions around models are going to emerge. We're experimenting with a set of abstractions, but these are going to emerge and they are going to be supported by both the programming model and then these kinds of innovations. LoRa's, activated LoRa's, I'm sure there are other innovations waiting that will allow us to write really complicated things.
37:57And if you step back and look at what's happening, things like what people call agents. Already agents contain some form of classical procedural imperative code. They always do. Sure. Right? And what people are figuring out is what should I do and what can the model figure out on its own. But that's a little bit of a contract on what I should do. Certain things today we are asking the model to do. We may be end or off just doing it outside because then you can still achieve a lot of the same functionality but get it done with a small model. So I think that's the journey we are on and that's what's exciting to us.
38:29That's why we think we are at a point where it's no longer we're not just doing AI. We actually have a new kind of computing element that we're going to know how to interact. And actually, that leads me to a question I wasn't going to get into to the end, but IBM's commitment to open sourcing these models. Meta has just asked, you know, people to help raise money for training future LLAMA models. And that's something I've kind of been used to ask Yan La Koon about. I mean, how long are they going to stay committed? this is such an expensive enterprise if you're just giving the models away. If you're doing it with capable but smaller models, you can do many more than Medican with LAM.
39:21Is that part of the strategy with these small models? Yeah, I think, look, our strategy was about small focused models from the get-go. And I think that was motivated less necessarily by whether that's the only sustainable way to put it in the open or not. We think that's the only sustainable way to get enterprises to at scale adopt AI. Because finally, they have to run these models. Yes, I could spend the capital and train the biggest model, but if that isn't necessarily that going to get deployed, and particularly with IBM strategy, we are looking for it to run across a hybrid environment. So we wanted to work on multiple clouds, on-prem infrastructure, diverse set of hardware.
40:00We've been very public about that too. run on NVIDIA, run on AMD hardware. We have our Spire chip where we run on our own mainframe. So hardware agnostic multiple environments. It's very important and that means that being able to optimize the model to get it running very very efficiently is really important. So if I deal with a huge object that isn't easy to manipulate, I'm basically shooting myself in the foot with kind of a strategy. So we were on small models anyway. So I don't think it was necessarily one way or informed by this and you can see that even in the philosophy I have taken, Granite 3.3 for example included a speech model, a vision model, a 2 billion better document processing model.
40:42Because eventually what is going to happen is you are going to solve a problem. To solve a problem or build an application you are going to have AI models work together. Some cases they are orchestrated by an agent, other cases you are stitching together. I have a document processing pipeline, I use a 2 billion parameter model. On the run time I use an 8 billion parameter model. If I have speech input, I use a speech model. I don't need one model to do all of these things. I can run these models differently. If that speech model needs to only run in a certain environment, I only need GPUs to run the speech model in that environment.
41:15If my document processing needs, I'm a hybrid cloud company, I'm using Watson X on Azure to run my document processing, that model needs to run on Azure. My application could be running on the mainframe, that model needs to run on Spire chip on the IBM mainframe. So this unique philosophy of fit for purpose decomposition we think is the right approach to build AI. So we've been on that path. Yes it does make it therefore much easier for us to iterate, rapidly release. Look at Granite 3. We released Granite 3 in October. 3.1 end of the year. 3.2 in January. 3.3 in February. We're going to be releasing Granite 4 by 2Q.
41:56That pace is possible because we have made this choice. And you can compare that to people who have made a different choice, whose release cycles tend to be six, nine months a year. It's just the way things are. Yeah. Yeah, I wanted to switch to more forward-looking, and not that this isn't forward-looking, but continuous learning in the search for... I mean right now although you've managed with these small models to train them at inference time, that doesn't adjust the weights in the underlying model. So there's no learning in the underlying model during inference. Correct. Inference scaling is on the top.
42:46Correct. the model remains the same. The only change is the LoRa. So you can drop in and take out LoRa. So that does give you runtime adaptability. Right, I see. Are you pursuing research? And I know that it's an active area of how you could update the weights of the underlying model during inference time to create a continuous learning loop. Right, so I think maybe to pass that a little bit. So I think first is the this notion of LoRa's and activated LoRa's already give us a way to to change notionally the weights of this composite thing instead of calling it a model. I want to call this as composite thing which consists of a base model and multiple adaptive LoRa's right.
43:36So already I get more flexibility than just having the one base model. I can turn it on and off. I can add a LoRa, I can remove a LoRa out of the system. So that already gives me modularity. So that's number one. Which means that I can deploy a system with five LoRa's. I'm collecting telemetry data of usage of the model. I can use that to go off and train another LoRa, plug it back into this model. So I don't necessarily have, it doesn't necessarily mean that every online learning I need to do has to go back and touch the base model. Now there are certain environments in which maybe I need to do that.
44:14So for example, I'm trying to collect chain of thought data or backtracking logic in an entirely new domain. I may need to seed some of that early on in the training of the base model, maybe in mid-training or something like that. So I may deploy a system, instrument it to understand how humans are doing that, collect the data, curate it, generate synthetic data and go back to it. That I think is a longer loop. But there are certain scenarios where I may be able to connect trajectory information like this. Just train a LoRa adapter that is trained on that and simply helps the model do that backtracking.
44:55Or I may be able to adapt a reward model that sits outside the base model that's trained on, Oh I see this is how humans backtrack in this domain. Let me get a reward model that's going to now teach the main model how to backtrack. And those reward models, because they are separate, I can update them and refine them on their own. So I think the way we are going to do online learning, until we figure out maybe a methodological innovation, is essentially by making the runtime that is actually serving AI highly modular. So reward models, LoRa's, I see as the mechanism to do that. Because every time you have to fall back on updating the base model, you're right, that's a long loop.
45:35Are there, have there been ideas around in real time updating the watch to the base model? Yes, but I haven't seen a practical deployable answer. Right now the base model is still a retraining activity. Right. And their small is easier to update, that's about it. I haven't seen anything fundamentally moving the needle. But I do think modularity gives you a good way to address that. Yeah, although what I'm talking about with continuous learning is that the model would increase its knowledge base without getting larger, maybe would have to get larger, I don't know. And so once you're working with it on inference, it gets smarter.
46:26I mean, this is, you know... The original, yeah, learn smarter. The hockey stick and intelligence, yeah. Correct. But I think what we are discussing is the mechanism for that to happen, right? So the long loop of that is let the system run, collect the data, then you come back and retrain the base model. That's the long loop. It's not automatic. It's the longest loop. The second faster loop is I can update some of these components quicker. And so I learn faster. I noticed that the model is not doing well for certain classes of queries. quickly collect the data, build an adapter that addresses that, build a better reward model.
47:01That's the second way to do that. I think there is a third way to do, which is at an application level. You deploy a model, look at where it's not doing well, and then use that data to train a second model and route. So I think at systems level, you can make the system more and more adaptive. I see. So I think those techniques are going to be very, very practical, because they allow you to react. Now, I do think that what you're asking for, which is this vision that the model discovers what it is not good at on its own versus a human being looking at it and deciding to train. I think the big piece there is this whole area of uncertainty quantification.
47:44So that's why we are working a lot on that. Because the first step to doing that is for the model to have a principled way to know what I'm good at, what I'm not good at. And if there is a region of space in this knowledge that I'm consistently having low uncertainty, then you can figure out what action to take. Today, all the actions are human driven. The human decides when to come back and retain Allura. You're talking about a system that automatically recognizes its faults and goes there. I still think that is still in the realm of innovation and research right now. But possible. But possible.
48:17But certainly possible. I view uncertainty and quantification as a significant element. Because that's the first step. You've got to know what you're not good at. If you continue to hallucinate your way and you don't know, then you don't know you have a problem. So uncertainty quantification, because remember that if the model produces, if a system produces an output, and the model does not get a judgment, it doesn't even know whether it did well or not. Now today, some other external feedback signal says, I'm getting customer complaints. I'm not resolving these tickets, whatever they are. And a human looks at it and comes back and fixes it.
48:51So for the model to learn, the model has to automatically know what it is not good at. And that's why I think uncertainty of quantification is actually a really, really important piece of the puzzle for self-improvement. Okay.
49:07Yeah, I'm just... Yeah, this was a question off of that. if inference time weight adjustment is possible, would it happen at layer-specific updates? How fine-grained would that weight adjustment be? Or would it be, I mean, the problem with right now with catastrophic forgetting is it may update overwrite previous knowledge. Right. But would it be done at a very granular level, like different nodes or different layers? I think the reason I keep going back to LoRa is that's actually a principle way to do that. Because if you draw the box not around the base model, But around the base model with the LoRa's, you are adjusting weights.
50:07But you're adjusting weights by moving from a certain computation into activating a LoRa, which does something different. It's sort of like inserting something at that point. So that's really what you're doing. So at a system level, LoRa's do give you an opportunity to update the effective network that is giving you the answer. That's different from... And the reason I think that is useful is, As you said, changing the weights of a base model is complicated because you took something that was created by a massive optimization, which is balancing many, many things. You're making some local changes whose implications you don't know.
50:49That's what leads to all sorts of challenges. The philosophy with lower eyes, let's decompose it to specific pieces of functionality. I want to teach the model a piece of knowledge. I put a knowledge lower eye. I want the model to be good at hallucination. So when you do it that way, you are isolating these things. And so you have less worry about conflicting with what was taught to the base model. You have individually optimized those LoRa's to do what they've done. And you're putting this together almost like a workflow. I think that's a more viable way to do what you're thinking about online learning.
51:24Because now it allows you to reason about this in a meaningful way. It allows you to separate functionality. So you know that if performance comes down when a LoRa was enabled, then you can go debug that LoRa. If I did this update of the base model, I don't even know how I would, let's say I did that, Debugability, traceability, all that becomes extremely high. So I keep going back to, I think the way we are going to get to online changes is really by not treating the AI system as one big network, But it's really a base network with lots of plugins. And we are probably early in the innovation associated with it.
52:07And so I think that's a more viable way in which we will get online learning. Otherwise, it becomes impossible to reason about it. These optimizations, because training is effectively a big optimization problem. Over all the data that you send, minimize this loss function. So it's very hard to reason about. Loras have the advantage that I optimize them for one thing and I plug it in. I trained another Laura, I optimize it for something else. I plug it in. It's much easier for me to reason about it, swap it in and out. So I feel like that, again this is just me speaking intuitively. Maybe there is a brainwave, a way to do this that I'm not thinking about.
52:44But I view this as a very pragmatic way to get the effect that you're thinking about. Yeah. Is the Granite series, is that the main thrust on generative AI right now? At IBM Research, it's bigger. So maybe I'll step back to talk about our overall focus, right? So one is certainly obviously the models themselves and everything we have talked about. A second big area for focus for us is on the data side. Yep. From an enterprise context, unstructured data is becoming important now because of Gen AI. So a huge focus for us is making our WatsonX.Data lake house really best of breed for unstructured data, which means thinking about all aspects of data management from ingestion to processing to ETL pipelines to security, access control, governance.
53:34So that's a big, big focus here with a lot of innovation from research. So you'll see, again, announcements around agentic rag, announcements around unstructured data processing pipelines and document understanding. That's driven by research innovation. That's a big thrust, which is how do you make enterprise data management ready for unstructured data. There's a second side to the data side, which is there's a lot of work in enterprises that data engineers or data stewards or people. And we call this broadly this term data ops, right? Yeah. Running, this is even with structured data. Forget unstructured data.
54:08Today, even before Gen AI, there's a lot of time spent in creating pipelines of unstructured data, cleaning them, morphing them, transforming them, moving it from one place to the other. So data ops is a very expensive piece. So one of the areas we're working on is how does agentic AI help those personas who are data stewards and data engineers do their job better, faster, so that's applying AI to data management. So on the data side, we sort of look at it as two right? Preparing data for Gen AI and then using Gen AI to make data management more efficient and effective. So that's the second thrust.
54:49The third thrust for us is everything to do with agents and agent middleware. So we believe where the world is going very much are just like our philosophy around models, right? Yeah. We are about fit for purpose models that you can pull together. Same for agents. We think the way the world is going to be built is multiple agents you have to orchestrate to get your work done. Because it's much easier to securely build an effective agent that does one thing as opposed to one humongous agent that does everything under the planet. Which means managing agents, the lifecycle of agents, governing agents, the security of agents, orchestration of agents, that's going to become an important focus.
55:29So we have from our product portfolio Watson X Orchestrate is our flagship product for agentic middleware. And a lot of what we doing from research is looking at problems like observability of agents. Because you have to look at agents in action just as you observe IT systems. How do you do agent ops and observability? How do you securely make tools available to agents? How do you look at security and authorization of agents? It's a big issue because today if you're going to have agents. So imagine Craig you're an IBM employee. I put in front of you an agent that lets you do HR tasks. Now that HR agent is going to be equipped with tools to talk to HR systems, make operations changes.
56:14Those tools have to only have your authorization. So that you are only allowed to do. So securely delegating authorizations and permissions is a big area that touches on the overall security fabric. So that's a big focus for a security governance and operational lifecycle of agents and agent orchestration. is an area. And we are building also agents in three specific domains. Those are where we have IBM businesses. So we have a big focus on agents for software development, but not general purpose software development. We have doubled down on areas where we have a big IBM business. We have a huge franchise on COBOL on the mainframe.
56:52So we have Watson Code Assistant for COBOL, Watson Code Assistant for Java. So that building agents for software development in those domains. A second big focus for us is building agents for IT automation. We've had a rich portfolio now of IT automation, combination of organic and inorganic acquisitions, right? So across our Aptio, Instana, Turbonomic portfolio, we address the entire lifecycle of IT, from cost management to operations to security and compliance. So we're building agents to help those personas, whether it's a CISO, whether it's an SRE, whether it's the IT analyst managing cost, agents to help their work do better.
57:32And then the third area we are building agents is for our Maximo portfolio, which is used to manage physical assets. So Maximo today is used to manage physical assets like in factories and utility companies. And so there, the equivalent of the SREs, there is equivalent of maintenance personnel who have to keep these systems up and running, look at reliability, manage faults, building agents. So those three domains, software engineering, IT automation, and physical asset management. We're building agents. And as we build those agents, we understand what agent middleware is needed. So that drives our innovation in middleware.
58:07So models, data, agents, and three domain-specific agents. That's basically the... I should probably leave it there, but this has been fascinating for me. Thank you. It's been a great conversation. Yeah, it's been a great conversation. Yeah. This episode is brought to you by Tasty Trade. On Eye on AI, we talk a lot about how artificial intelligence is changing how people analyze information, spot patterns, and make more informed decisions. Markets are no different. The edge increasingly comes from having the right tools, the right data, and the ability to understand risk clearly. That's one of the reasons I like what Tasty Trade is building.
58:51With Tasty Trade, you can trade stocks, options, futures, and crypto all in one platform with low commissions, including zero commissions on stocks and crypto. so you keep more of what you earn. The platform is packed with advanced charting tools, backtesting, strategy selection, and risk analysis tools that help you think in probabilities rather than guesses. They've also introduced an AI-powered search feature that can help you discover symbols aligned with your interests, which is a smart way to explore markets more intentionally. For active traders, there are tools like Active Trader Mode, One-Click Trading, and Smart Order Tracking.
59:42And if you're still learning, Tasty Trade offers dozens of free educational courses, plus live support from their trade desk reps during trading hours. If you're serious about trading in a world increasingly shaped by technology, check out Tasty Trade. Visit tastytrade.com to start your trading journey today. I'm going to myself. Tasty Trade Inc. is a registered broker-dealer and member of FINLA, NFA, and SIPC.
From the publisher
Why IBM Is Betting Everything on Small AI Models
In this episode of Eye on AI, Craig Smith sits down with Sriram Raghavan, Vice President of AI at IBM Research, to explore one of the most important debates in enterprise AI right now. Do you actually need a massive model to get world class results? IBM's answer is no, and Sriram breaks down exactly why.
Sriram explains why IBM chose to train its Granite models directly using reinforcement learning rather than distilling from larger models like most of the industry. The reason goes beyond performance. It comes down to data lineage, safety alignment, and a belief that small, efficient models are the only sustainable path for enterprises running AI across hybrid cloud environments.
We get into the full technical stack behind that bet. How data quality has replaced model size as the real competitive advantage. Why parameter count is becoming the wrong metric entirely. How IBM's inference time scaling techniques allow an 8 billion parameter model to match the performance of GPT-4o and Claude 3.5 on code and math benchmarks. And why IBM is pioneering a new concept called Generative Computing, which treats AI models not as prompt receivers but as programmable computing elements with runtimes, modular LoRA adapters, and proper programming abstractions.
Sriram also shares where IBM Research is headed next, including breakthroughs in continuous learning, agent orchestration, and making unstructured enterprise data actually usable at scale.
Subscribe for more conversations with the people building the future of AI and emerging technology.
Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI
(00:00) Why IBM Skips Distillation and Trains Small Models Directly
(04:50) Did We Even Need Giant AI Models in the First Place?
(08:12) How Data Quality Became the New Competitive Moat
(11:54) Why Parameter Count Is the Wrong Way to Measure a Model
(15:36) Reinforcement Learning Without Losing Broad Capabilities
(22:05) Inference Time Scaling: Getting Big Model Results From Small Models
(28:12) Generative Computing: Treating AI as a Programming Element
(36:40) Why IBM Open Sources and How Small Models Make It Sustainable
(41:25) The Path to Continuous Learning Without Rewriting Weights
(51:00) IBM's Full Roadmap: Models, Data, and Agents




