In short
Why India has been slow to adopt AI and whether it can catch up; IBM’s enterprise-focused AI research in hybrid cloud, generative AI, agentic middleware, and continual learning.
Guest
Amith Singhee, Director of IBM Research in India and CTO role for IBM’s business in India. Background: electrical engineering at IIT Kharagpur; MS and PhD at Carnegie Mellon (electronic design automation); 18 years at IBM (8 years at Yorktown Heights); later work in AI/data analytics, IoT/retail/fashion, then hybrid cloud after IBM acquired Red Hat (2019–2020). Focus now: generative AI for enterprises, Granite open models, WatsonX code assistants, time series/geospatial (satellite imagery), and quantum software.
Key claims
India can catch up if “meaningful investment + talent + infrastructure” align (data centers/GPUs). IBM develops portable AI for regulated clients via hybrid cloud. Customization must avoid catastrophic forgetting; continual learning is needed for frequent updates.
Notable examples
weather/grid simulation with DTE Energy and PG&E; COBOL/mainframe code assistance in WatsonX (updated every 4–6 weeks); satellite-image models; IBM collaborations with IITs and the AI Alliance.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIndia's AI Investment and Infrastructure
0:45 to 3:00
Discussion on India's investment in AI and necessary infrastructure for growth.
“I'm a director of IBM research in India, and I also play a CTO role for the business here in terms of sharing our technology vision with the ecosystem and engaging on that front.”
Amit Singhee's Background and Role
3:00 to 5:45
Amit Singhee shares his professional background and role at IBM.
“So we looked more at IoT and retail and fashion and how deep learning can make a difference there.”
IBM's AI Innovations and Research
5:45 to 8:00
Exploration of IBM's focus on AI, hybrid cloud, and generative AI.
“ecosystem, whether it's the government, startups, companies, developers, again, from an innovation roadmap, let's.”
Talent Development and Innovation in India
8:00 to 10:00
Insights on attracting talent and the role of research in India’s AI landscape.
“It also helps us engage with the market and understand variations of the problems that are relevant, let's say, in a developing country like India.”
Comparing AI Ecosystems: India and China
10:00 to 12:15
Discussion on the differences in AI development between India and China.
“It was the sudden transformation to generative AI that allowed them to catch up because they had brought back these people, But they were far behind.”
Collaboration Between Industry and Academia
12:15 to 14:07
How IBM collaborates with academic institutions in India.
“So here, India does have this notion of professor of practice.”
Collaborations with Indian Academic Institutions
14:07 to 14:45
Learn how IBM collaborates with Indian universities for AI research.
“But if somebody really wants to do at scale experiments, then it's not that easy to do.”
AI Alliance and Open Collaboration
14:45 to 15:40
Discover the AI Alliance and its collaborative efforts in India.
“would be four or five different projects running at a given point of time with co-PIs.”
Global Research Collaboration at IBM
15:40 to 16:27
Explore how IBM conducts research across various global locations.
“And in India, we've had some academic institutions like IIT Bombay, IIT Madras, IIT Jodhpur, some startups, AI, and some nonprofits and some large system integrators become members.”
Challenges for AI Development in India
16:27 to 18:34
Understand the challenges India faces in becoming a major AI player.
“So you may be working on hybrid cloud solutions, and we can talk about the research, but you would be working with a team in the U.S.”
Show all 25 chapters
Optimism for India’s AI Future
18:34 to 20:09
Hear about the potential for India to catch up in AI development.
“What's your view of why India has been slow?”
Investment and Infrastructure in AI
20:09 to 21:44
Learn about the critical ingredients for AI growth in India.
“hoping to speak to a startup person in Chennai because that has been another issue is the startup ecosystem and all the barriers to a startup.”
Hybrid Cloud and AI Research Focus
21:44 to 23:14
Explore how hybrid cloud solutions enable AI research at IBM.
“But the infrastructure is still in a ramp up stage.”
Global Relevance of AI Models
23:14 to 24:35
Understand the global approach to developing AI models at IBM.
“But if I have a hybrid cloud mindset, then I might say, look, I don't want to develop one Gemini on one place.”
Advantages of Diverse Data in AI Training
24:35 to 25:56
Learn how diverse data enhances model training and generalization.
“And until now, we have supported mostly Indian languages have not been primarily supported.”
Customizing AI Models for Specific Needs
25:56 to 28:07
Discover how banks and businesses can customize AI models effectively.
“Is that because of the data that you're using from India?”
Customizing AI Models for Business Needs
28:07 to 29:10
Learn about the techniques for fine-tuning AI models to meet specific business requirements.
“But when somebody, if a bank wants to do something specific with that model, There, the model is already trained.”
Challenges of Data and Model Training
29:10 to 31:05
Explore the challenges of maintaining general skills in AI models while fine-tuning with limited data.
“For example, there's data mixing strategies.”
Continual Learning and Catastrophic Forgetting
31:05 to 33:53
Understand the concepts of catastrophic forgetting and how continual learning can enhance AI models.
“So if it's us doing that for the client and it's the granite model, then of course we have the data.”
Real-World Applications in COBOL and Legacy Systems
33:53 to 36:29
Discover how AI models are applied to modernize and maintain legacy COBOL systems.
“rather than keeping all of the control with the model provider.”
Integrating Business Context with AI Models
36:29 to 39:22
Learn about the importance of business context in AI models for effective code interpretation.
“So it's already in practice, but I think solving it more generally so that you don't need PhDs, you just need a button and it should just work.”
Managing AI Agents and Their Security
39:22 to 42:05
Examine the orchestration and security measures required for AI agents in business processes.
“And you mentioned a middleware for agents.”
The Hype vs. Reality of AI Agents
42:06 to 43:23
Explore the gap between the hype surrounding AI agents and their practical applications in enterprises.
“Mike, I spoke to Kumar, I think her name was at Microsoft, about she was talking about societies of agents that will be working together invisibly doing a lot of the routine stuff has to happen in business and society.”
Challenges in Enterprise AI Integration
43:23 to 44:55
Understand the engineering challenges enterprises face in integrating AI with existing business processes.
“Yeah, I think the hype is moving faster than it should.”
Preparing for an AI-Centric Future
44:55 to 46:51
Learn what skills and knowledge young engineers in India should focus on to thrive in an AI-driven job market.
“there's very high unemployment among IAT grads.”
Transcript
Automatic transcript. May contain errors.0:00What's your view of why India has been slow? And then do you think it's possible for India to catch up? What I have seen is India has taken more time to bring in the investment and the intent at the level needed that some of the other countries have done. It is a perfect storm of meaningful investment coming together with the talent and with infrastructure in a single entity. So once all these three things come together, then it's going to move really fast. I don't know if we can become number one or number two in the next two years, but if we look at it from the lens of are we maximally capable of applying the latest AI tech for the benefit of India on our own, I think we can get there.
0:40Whether we can make the next three trillion dollar startups, that's a whole other question. I'm Amit Singhi. I'm a director of IBM research in India, and I also play a CTO role for the business here in terms of sharing our technology vision with the ecosystem and engaging on that front. I've been in IBM almost 18 years now. My background is in engineering. I studied electrical engineering at the Indian Institute of Technology in Kharagpur, IIT Kharagpur, after which I went to the U.S., did my master's at Carnegie Mellon. This was in a field called electronic design automation. So I switched a little bit into the semiconductor and microelectronics area.
1:30I worked for a couple of years and went back to my Ph.D. back at CMU and then joined IBM Research in Yorktown Heights. Oh, really? Yes. I spent almost eight years there. Beautiful building. Where did you live? I used to live in Yonkers. Oh, yeah. Because my wife used to work in the city, so it was convenient. Yeah, and I started off working in semiconductors from the design automation lens, which was basically creating software tools for chip designers. Right. So it was at the intersection of computer science, electrical engineering, math. And then it's been 18 years in IBM. The last 10 have been in India.
2:15I've moved around in terms of teams and areas over the years, moved into more AI and applications and smarter energy. This was in the 2010 to 2015 time range. That was the period of big data and analytics and lots of utility companies were looking to use that because they're data rich. And we did some great projects at that time internally and also with utilities like DTE Energy in Michigan and PG &E in California in trying to simulate the weather and its impact on the grid. Once I moved back to India, I was more focused on things where we could leverage collaborations in the market here. So we looked more at IoT and retail and fashion and how deep learning can make a difference there.
3:10And then in around 2019, 2020, when we had our new CEO, we acquired Red Hat and IBM really pivoted into this hybrid cloud strategy. Within the research division also, we saw an opportunity to double down and help the company with innovations. So I started an effort around hybrid cloud research. out of the India team, working with our global research team, and looked at lots of interesting problems at the intersection of hybrid cloud and AI. And then I got into this role about three years ago as director. And now we are heavily focused on generative AI in the lab here. Not a surprise. But we look at all the layers of the stack and more from an enterprise lens rather than a consumer lens, which means we look at the stack in terms of enabling hybrid cloud deployments.
4:03So you should be able to deploy the models and the AI workloads wherever you want. If you're a bank and you have a highly regulated enterprise, you might want to decide where the AI goes. So we look at optimizing the software stock for AI. We also work on training and developing our Granite model series, which is in the open. And we have a lot of work also on building agentic middleware, which allows enterprises to build agentic solutions and also in the actual solutions themselves, specifically in areas like software development. IBM has this WatsonX code assistant series of products. We do a lot of the AI research that goes into those.
4:49Also in the IT automation area, business automation, data. How do you use AI to better manage enterprise data? And we have some very interesting angles to the AI work we do. For example, in the time series domain and geospatial domain, satellite imagery, where can you have, instead of a language model, a model that understands satellite images. So good spread of work. We also have a growing team in quantum computing, especially at the software layer. And we are very engaged with the ecosystem here. also that's one of the unique things about the research team in IBM in India is we've been here 27 years we started off within IIT Delhi we have very strong ties to many academic institutions they see us as collaborators because of research ethos and we're also very engaged with other parts of the ecosystem, whether it's the government, startups, companies, developers, again, from an innovation roadmap, let's.
5:57Yeah. As I said, I'm interested in where India is generally in the AI journey. And I had a long conversation with Amishik Singh in Delhi about why the government was late in getting on board or in forming a vision, but they are very focused now and there's a lot of ambition. Why do companies like IBM, why did IBM set up an idea? As I said, is it a market strategy or is it a human resources strategy? Because there are a lot of talented engineers here or is it built around some specific research that was developed in India and for which India is has the expertise? How do you see it? So we IBM came to India well before the AI revolution right so even the research team's been here for 27 years and the company's been here longer.
7:12So the beginnings were around being a global company, seeing the talent potential of a growing economy in the late 90s. It made sense to have a presence here. It's worked out well, because that potential is realizing and definitely the talent that we get access to here adds to talent we have elsewhere. So it's a very valid pool of deep tech talent with global mindset that helps us as a global team. Just us in IBM Research, the projects we do are not isolated by geography. It's part of a global strategy. And so having that combination of hard skills and soft skills available here does help our global mission.
8:06It also helps us engage with the market and understand variations of the problems that are relevant, let's say, in a developing country like India. So then the technology we develop is more generalized and more applicable. And then when it lands up in the products, again, then more relevant. And it's not just India. India represents a lot of the global south. So when you look at the problems from this lens and bring it into our mainstream work, it makes the solutions relevant to many other countries. So it's a mix of both, I'd say, the talent and the understanding of the market. Talent is surely a big driver.
8:46Yeah. And on the talent, why did you come back? I because again I had this long conversation with Malsam, Professor Malsam, about the you about the difficulty of getting people to come back after their PhDs because there's so many opportunities in the United States so yeah for me it was simple and personal I actually spent 15 years in the US and I loved it I could have stayed there forever but I really wanted me to be back close to family. Sure. And I felt, okay, the time in life with our family and everything was just right. And we moved back mainly for that reason. Yeah. As I said, I spent a lot of my life in China and Maelson talked about this, but also I lived through it in China.
9:41There was a moment when China's economy had reached a level that they had tremendous foreign reserves and they very deliberately started offering competitive salaries to bring expertise, research expertise back. But it wasn't that alone. It was the sudden transformation to generative AI that allowed them to catch up because they had brought back these people, But they were far behind. But once generative AI, the transformer algorithm, they suddenly didn't leapfrog the whole supervised learning optimization phase. And now they're competitors. So it looks to me that India has the opportunity to do the same thing.
10:43Are you guys able to attract talent from the Indian talent, Indian national talent back from the United States? So speaking from the lens of IBM Research, we actually find our talent in the local market mostly. And we do have people like me who move back to the lab and join IBM Research. but mostly it is because they're moving back for their own reasons rather than us convincing them. But this is from the lens of our lab, and we have eminence in the country in the research space. So if there is good talent, there's a high chance that they'd want to go to a research lab like ours. It's definitely not the case at large that there's a lot of local talent at that level available at scale in the way it would be in the U.S.
11:38And it is something that does need more effort. Yeah. Again, and are you hiring from the IITs? We hire from the IITs. We also hire from some of the other colleges. We hire PhDs. We hire masters. We also hire undergrad, strong undergrads. And does IBM have a policy of, actually, I don't know if Parta is still teaching. Do you? I don't know. Yeah. But he's now full time at DeepMind. If you hire somebody from a PhD student from the IITs, can they or a professor, can they continue to teach at the IITs or, you know, as they do in the U.S.? A lot of people have dual roles. So here, India does have this notion of professor of practice.
12:32and in some cases if somebody is really passionate to do that there are mechanisms to put that in place but what we see is a lot of people just want to if they are coming into an industrial research lab they want to be successful there and make a difference and that tends to take up a lot of your time and energy but yeah there's a professor of practice mechanism that's available Yeah, and the resources, you were saying that you collaborate a lot with the IETs. One of the reasons in the U.S. that people go into industries is because of the resources, the compute and data resources that aren't available to academic institutions.
13:21Is that the same here? That's one of the drivers that pulls people toward industry? It is, especially in this generative AI era. It's not, yeah, there's not a huge amount of resources available to academics in general. Some of them would have gotten grants, but again, that's very specific cases. Whereas in industrial settings, sometimes they can get more access. So that depends what drives the faculty. If they, a lot of them do great work, which doesn't need a lot of GPUs because they're doing theoretical things or they're trying to figure out nifty ways of improving something. But if somebody really wants to do at scale experiments, then it's not that easy to do.
14:11Yeah. And in the collaborations that you have with academic institutions in India, how does that work? You provide compute or you? So the way it works is there's a couple of different ways. We have long-term research collaborations contractual with IIT Bombay, IIT Delhi, IISC, where there would be four or five different projects running at a given point of time with co-PIs. A faculty member and one of our researchers or two of our researchers would be co-PIs directing the work. So the student gets the benefit of industrial mentoring and practical problems. And faculty groups, they benefit from steering research in areas where industry is focused on and creates a job pipeline for some of these deep skills.
15:13And these would be funded programs by IBM. The other way also we work is in the open. There's this thing called the AI Alliance that was started two years ago. It's a community of organizations globally that have said, okay, we want to do some work in the open to build AI, to create data sets, to set up responsible practices for AI. IBM is very active in that. And in India, we've had some academic institutions like IIT Bombay, IIT Madras, IIT Jodhpur, some startups, AI, and some nonprofits and some large system integrators become members. And so there, these members might find common interests and they collaborate in the open.
16:06Each one gives something. Somebody gives people, somebody gives a few GPUs or data sets, and we work on things. And that's a global alliance. That's a global alliance, yes. Yeah, that was another question. So you were saying that research is not geographic at IBM. Right. So you may be working on hybrid cloud solutions, and we can talk about the research, but you would be working with a team in the U.S. or with a team in France or wherever. Yeah. Right now, there's so much balkanization of the world politically, unfortunately. Are you able, does that interfere at all with like working with IBM in China or working with IBM in, I don't know if IBM has a big research presence in Russia, but I'm guessing they might.
17:14We don't have research presence in China or Russia. So we have research presence in Japan, in India, in Israel, in the UK, in Switzerland, in the US and in parts of Africa. Of course, we have to be very clear and careful about all of our export compliance requirements. And we are all aware of that and we keep up to date with it. But given our footprint, we don't really run into the footprint that I spoke to you about. We don't run into a lot of issues typically. Wherever there's specialized technology that even within this group needs to be contained, we'll be careful about that. But that's limited.
18:02Yeah. And how do you feel about India on its, as I said at the beginning, what interested me about India is looking at it from China. and not understanding why India is not a player, at least yet, in the AI competition. I know researchers don't like to look at it as competition, but the political leaders certainly do. What's your view of why India has been slow? And then do you think it's possible for India to catch up as the India AI mission's ambition is? Yeah, so I don't know why things have happened as they happen, because it's up to decision makers in the government and what they've done.
19:02But what I have seen is India has taken more time to bring in the investment and the intent at the level needed that some of the other countries have done. And perhaps that was, that's politics. I don't know what it is, right? But definitely since last year, I would say, it's grown up several notches. and I'm very optimistic being engaged with many of these stakeholders and the intention that's there and the talent that's there. I think we have some ingredients if we can use those right. We can be quite relevant. I don't know if we can become number one or number two in the next two years, but if we look at it from the lens of are we maximally capable of applying the latest AI tech for the benefit of India on our own, I think we can get there.
20:02Whether we can make the next three trillion dollar startups, that's a whole other question. Yeah, as a matter of fact, I'm hoping to speak to a startup person in Chennai because that has been another issue is the startup ecosystem and all the barriers to a startup. The level of investment I think it's a bird of 1.2 or 1.6 billion. I can't remember the exact number. And a lot of that's going into data centers. and India is, I think, still only one gigawatt in consolidated data center power. What do you think is the critical ingredient? Because India has, with the IETs and the JEE, you've got the top talent coming out of the IETs.
21:12Yeah. And there's also an English language facility, which I would think helps. So what do you see as the biggest bottleneck? I think it is the perfect storm of meaningful investment coming together with the talent and with infrastructure in a single entity, whether it's a startup or an academic effort or whatever it is. And that combination hadn't happened until last year, for sure. Now, this year with India AI mission and some of the startups that are getting the major funding, I think it will happen. But the infrastructure is still in a ramp up stage. The money is there. People are there. Modulo, money getting released and all of that.
22:02But it's been committed. But I think our data centers have been building up. I don't know all the details but what I understand is getting all the GPUs alive and in a single data center so that you have 4 ,000 GPUs or whatever you need in a single place available for six months it's getting there so once all these three things come together then it's going to move really fast because I know that the skills and the knowledge of what needs to be done in terms of how do you train these models how do you prepare those skills are there these startups, our folks, people have figured it out. So they don't need to rediscover that.
22:44It's just that those three things have to come together in a sustained way so that people can do the work. And by infrastructure, you mean primarily data centers? Yes. Yeah. Okay, so the research that you're doing here, it's on hybrid cloud. Is that the main focus? The main focus is AI. But enabling AI with a hybrid cloud architecture. Right. Okay. So what that means is if I didn't have a hybrid cloud mindset, then I might just develop like a Gemini that runs on my cloud and I'm done. But if I have a hybrid cloud mindset, then I might say, look, I don't want to develop one Gemini on one place.
23:30I want to enable clients to use any model they want anywhere they want. and I want to develop models that are portable. I want to develop software technology that runs those models, uses those models in a way so that it's portable. So it's still AI research, but it's architected and prioritized so that it can be deployed anywhere, on any cloud, on-premise data centers, on power systems, on mainframe computers and things like that. So that's the lens we take. But the core research we do is, again, it's LLMs and image models and agentic systems. So it's all the same fundamental technology. And the LLMs that you're working on here, are they English local language specific?
24:25And do they cover a range of languages? So, no. Again, the LLM work we do for our granite models, it's a global effort. Right. So it's just one version of the model that anybody can use. And until now, we have supported mostly Indian languages have not been primarily supported. But our team here has been working to add better support for Indian languages into the same model series. So the next version that's going to come out, we expect it to be a lot more capable. The other thing we are also doing is as part of our collaborations in the AI Alliance and other areas where somebody else might be developing India-specific models, we work with them to show how those models could be made hybrid cloud compliant or how could you scale them better.
25:18Because we don't see IBM really building India-specific models or Germany-specific models. It would be one common model which does reasonable everywhere. But it's likely someone will come up with a model that's really tuned for that market. And so we also want to work with them to allow them to scale and deploy their models wherever clients might want. Yeah. One thing you said earlier that I thought was interesting is that by working in India on these global efforts, it helps generalize models. Is that because of the data that you're using from India? So I'll give you a couple of examples. If you are only looking at data that's in the West, you'll have a lot of data on the internet.
26:14But if you say, okay, let me look at digitized data in India, that's way less. It doesn't reach the levels. then you're forced to think how do I train models when I have less data and still get high quality that helps you generalize your approach so you're not training models with the assumption that I'll always have a huge amount of data and that's useful because when we go to a bank or an enterprise and they say look I want to customize this model for my purposes but I only have x amount of data can you do it. So that generalized approach of being able to work with smaller amounts of data can be useful in lots of these settings.
26:56That's one example. Another example is in these transformers, there's this notion of tokenization where you break up the words into little tokens like syllables. Now, which tokens you use can get biased based on which languages you've looked at. If you've not looked at Hindi and Indian languages carefully, then when you try and apply the models in India, those syllables are not going to be optimal and the model might not understand or interpret that language right. But if you have done that, then you can generalize your tokenization approach. These are some of the things that happen on the model side.
27:36Yeah, that's interesting. And this is something, a general question that I've never quite understood. But when you say you're working on the granite series, so those are open weight models available to anybody. the foundational training is what you're saying is that you contribute to that foundational training data so that the model will generalize across cultures and continents. Yeah. But when somebody, if a bank wants to do something specific with that model, There, the model is already trained. You're talking about fine-tuning, extended training. Yeah, there's different techniques, but generally, we just call it customization.
28:35Right. So you make the model even better for that business. Yeah. And how much data is, obviously, it depends what you're trying to optimize for. But you need much less data to do that.
28:53But when you were talking about helping customize if a customer has only a little bit of data, what do you do? Where is the research leading? Yeah, so there are things like there's a couple of tricks, right? For example, there's data mixing strategies. case. So if you have X amount of data, right, which is small, and I take a nice granite model and I tune it on that, what could happen is it might get really good on the tests you run, but it'll really lose all of its general capabilities that it had. It overfits. So it forgets things. Now that might become a problem when you deploy it out there in the field because you only test on 100 examples, but your consumers are going to throw all kinds of questions and things at it.
29:51Relevant to your business, but it does need that general knowledge and general skill there too. So you need to then figure out how do you maintain that general skill while infusing your business-specific skills in this tuned model. And there's tricks around how do you mix the old data with the new data and what proportion do you use? how do you feed it to the model? There's this notion of curriculum training, like how you might train someone who joined New York Times. You might ease them into it. So all of these recipes, we call them recipes, have to be tuned to the data scenario you have in that client's environment.
30:33There's also other tricks around synthetic data generation. So you might say, okay, if you only got X data, I will have some clever tools that will understand the distribution of that data and synthesize a lot of data that looks like that. It's not supernatural, but it's good enough so that if the model sees the realistic data and sees some synthetic versions also, that's better than just showing it that little bit of it. So there's all these innovations around data that come in. And the data mixing, are you pulling data out of the original training data and mixing it with the fine tuning? If you have access to it.
31:14Yes. So if it's us doing that for the client and it's the granite model, then of course we have the data. But if it's somebody else trying to do it, they may not have access to the old data. Then there's other tricks you have to try where you try to make the model generate old data that it might remember and then use that back. Yeah, that's fascinating because this is a little bit off tangent. The big blocker in my mind to right now, we seem to have reached a plateau in generate AI models. or there's some incremental improvement. It feels like where we were with supervised learning, where the research community focused on optimization and getting smaller and smaller improvements or refinements, which are important for productization, but in terms of getting to ASI, it's not going to happen.
32:22So to me, the big blockers are catastrophic forgetting and continual learning, lifelong learning, whatever you call it. Does this research what you're talking about, it sounds like that might be relevant for solving? Yes, I wasn't using those terms, but it's exactly what you said, catastrophic forgetting and continual. Actually, catastrophic forgetting is no longer catastrophic. it's just forgetting because we've gotten past those early days where researchers would just do like bad data mixes and it'll just forget everything but now it's more like oh it used to be it used to get 100 marks on this kind of a task now it's getting 80 it doesn't go wrong but we still want to keep it at 90 plus because we know that's a useful skill in the wild world out there and you don't want to forget even that little bit.
33:22And continual learning is another important one because model versions change, new data comes along, you want to add new use cases. As a bank, you say, okay, initially I thought I'll train this, but then a year later you've got sort of new things and you don't want to go back to the scratch pad and do everything from the beginning. So these are all very valid problem statements that come up when we put this hybrid cloud and enterprise lens, where we want to give control to the client rather than keeping all of the control with the model provider. Does in lifelong or continual learning, if you can solve catastrophic forgetting, does the model do the parameters need to expand with as the data is added?
34:22So we try not to. Typically the model, as of now, the approaches that are being pushed and improved keep the same model architecture, the same parameters, same everything. It's just that you figure out new values for those parameters, either through extended training process or some reinforcement learning process or using these things called LoRa adapters, low rank adapters. But those low rank adapters, while they are additional parameters, but they overlay onto the same parameters at the end of the day. So, yeah, as of now, there's not, I am not aware of much that changes the model. Are you hopeful that the catastrophic forgetting will be solved to a degree that continual learning can become a reality?
Read the full transcript
35:16I think it can become a reality, yes. But I don't think it's going to be perfect. But I think it's going to become useful. Yeah. It'll be good enough to be useful. In fact, when we work on our products, like the WatsonX, we have this code assistant for COBOL, like legacy languages, which actually a lot of clients use. A lot of credit card transactions, bank transactions are actually handled on the mainframe. And those are all running on code written in COBOL, assembler, all these things. So none of the models that you find in the public domain do well on these because they haven't seen that kind of data.
36:01So we take our Granite models and then we customize them with our in-house expertise in IBM to become better on understanding COBOL, writing COBOL, converting COBOL to Java. and we release an updated version every four to six weeks, which means we've had to figure out continual learning and some aspects of continual learning for that. So it's already in practice, but I think solving it more generally so that you don't need PhDs, you just need a button and it should just work. I think we're not there yet. Yeah, yeah. And that's interesting. So that's because there is this massive installed base of cobalt.
36:50Yes. And there have been a lot of, there's cottage industry in translating or modernizing cobalt, but because to find an engineer that understands the language is getting harder and harder. So these models are taking over that role. They can understand the language. They're not necessarily translating it into a more modern language, but they allow engineers to work with the companies. It's all of those. So you can use it to explain some COBOL that's looking cryptic to you. You can say, explain this to me in English. What does it do? You can use it to write new COBOL. Maybe you've got like a complex tax calculation code and you're like, oh, the tax rate changed.
37:50Where should I go make the change? So the model will try to guide you. You can use it also to convert some COBOL code to something like Java because it understands COBOL and Java. And so it would read the COBOL and say, hey, this is a Java that you could use. You still have to double check it, but it can actually translate COBOL to some more model languages. The new thing that we are trying to do on is not just the model, but the system so that it can bring in business context. Because this COBOL code that clients have, it actually implements their business process. How is that line of state tax calculated on my mobile bill?
38:35Behind that, there's some business rules. That's what's coded up. And if you don't know COBOL and if you didn't know 40 years ago, why was this written in this way? That knowledge of those business rules is getting lost or it's locked in, I would say, right? So that's a huge risk for our clients because if let's say 10 years from now, they don't have anybody who understands why that tax calculation was done that way or what is actually happening. It's hard for them to maintain it or there's a new regulation that requires a change. It's a huge risk. So interpreting that entire code base and creating a business description of that is one very interesting problem we're working on.
39:21Yeah. And you mentioned a middleware for agents. Is that an orchestration layer to manage multiple agents or the coordination or cooperation of agents? There is that. That's the most obvious one. There's a product we have called WatsonX Orchestrate that does that. But then that's not all. What happens is, let's say you use a powerful LLM as an agent where you say, book flight tickets for me or promote somebody or whatever. Now, nowadays, these agents, they first make a plan dynamically. To do this, I've got to do this, and then that will call that tool, call that API, read this file, and then I'll do that.
40:08And then based on that, it'll take actions. It'll call some API. It'll generate the code for calling an API, or it'll open up a file on the disk and read its contents and then decide what to do. So these things are called tools for the LLM. And if you don't provide good tool descriptions to the LLM, it may not know what it actually does. Not so different compared to humans. If you put a new employee into a complex business process, but you don't give them very good standard operating procedure documents and you just give high level descriptions, they're going to make mistakes and do something.
40:50So that's one example of now. So now you're shifting the who's going to write those descriptions. Now, again, you need humans, right? So you can create tools that actually generate good descriptions or that they test. Are these tools actually doing what the description says it does or did you miss something? So there's a lot of that's one example. Another is security. If you've got an agent that is performing some business task in the SAP system or something on your behalf, you need the same identity authorization, auditability, and all that you would need if you were doing it. But now you're offloading it to this AI system, and there might be thousands of users using the same system.
41:39So you need security mechanisms to track that this request came from person X, their enterprise authorization is Y. And so I need to like that particular call needs to. So there's all of this machinery you need, which is not just the orchestration. So we as an IBM research, not just the India team, we are working on a lot of these things so that then we have more mature capabilities. And are you seeing there, Mike, I spoke to Kumar, I think her name was at Microsoft, about she was talking about societies of agents that will be working together invisibly doing a lot of the routine stuff has to happen in business and society.
42:33but there's been very slow uptake by the enterprise uh it are and largely because it's a trust issue not just security but um and i played around with agents uh and you set it off to do something and you go away and come back and realize for it's been waiting for 45 minutes for you to say yes.
43:06As IBM, of course, you believe in the promise of agents, but it just seems that there's a lot of hype about agents, but practical use is lagging, and it's lagging because of these reliability issues. Yeah, I think the hype is moving faster than it should. As it always does. As it always does. So, again, for enterprises, they want to be sure it's going to work in the way things used to work when I didn't have the agent in there. and just because I can do things with chat GPT using flight tickets doesn't mean that this thing is going to integrate right with my SAP systems and with my workday and with and it'll have the right user experience because ultimately some end user is going to work with it and it needs to fit into that business process.
44:11So all of these engineering things are still left to be done. I'm not saying these are rocket science, but they need time and effort. They need a lot of design and iteration to figure out. And alongside that, there's also this thing that not everybody is going to integrate a public frontier model into their internal business process, right? And so when you have an on-prem deployment of some model, getting it to do what OpenAI figured out is not trivial. So that's the gap, this sort of engineering and sort of enterprisification that will take a bit of time. It will happen, but it'll just have to play out.
44:54For young engineers in India, there's very high unemployment among IAT grads. And I get asked this question, I'm a journalist, I have no idea, but people ask me all the time, what should I study? What would be your answer? What should an 18-year-old study today to be part of the AI future that we're looking at? Yeah, so if it's an AI-centric future, because they should study things that are not AI also, because they're still a mechanical engineer. So all those fields need. Although AI will take care of a lot of that. Yeah, but you do need the domain expertise. If you don't understand mechanical engineering, even if you're using AI, you will not be useful to an automotive company.
45:50So I would say two things are important. Understanding the fundamentals is still important, even if you're going to use AI in the future. because people hiring you will always hire the person who understands the fundamentals and can use AI versus someone who doesn't and can use AI. That's important. That'll empower you. And the other is, I think, we're getting beyond this area of I've learned a skill, I'm good in that, and I'll have a career in it. That's gone, right? So the more diverse experiences you have so that you condition yourself to continuous learning. Also, humans need to have continual learning and application of that learning very quickly.
46:37That's a skill that is important. I think it's not just on the students. Academic institutions have to create that environment. That's the combination at this point, at least for the next few years. Okay.
From the publisher
What if the country that trains the world's engineers finally built the infrastructure to match its talent?
In this episode of Eye on AI, Craig Smith sits down with Amith Singhee, Director of IBM Research India and CTO of IBM India and South Asia, to explore where India actually stands in the global AI race and what it will take to close the gap.
Amith gives an honest, ground-level assessment of why India has been slow to compete. The talent has always been there. But until recently, the investment, the compute infrastructure, and the institutional intent hadn't come together in a sustained, coordinated way. That's changing, and Amith explains exactly what's different now.
He walks through IBM Research India's 27-year presence in the country, the research it's doing on foundation models, hybrid cloud AI deployment, agentic systems, and quantum computing. He also explains why building AI from India doesn't just help India. Working with less data, less compute, and more linguistic diversity forces better engineering and makes IBM's models more generalizable for the entire world.
We also get deep into the technical frontier. Why catastrophic forgetting is one of the key unsolved problems standing between current AI and anything more capable. How IBM is already shipping continual learning in practice through its COBOL modernization tools, helping enterprises decode decades of legacy code before the engineers who wrote it are gone. And why agentic AI, for all the hype, still has a mountain of unglamorous enterprise engineering left to climb before it becomes truly reliable.
Plus, what Amith would tell an 18-year-old engineer in India today about what skills will actually matter in an AI-driven world.
Subscribe for more conversations with the people shaping the future of AI and emerging technology.
Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI
(00:00) Introduction and Amith Singhee's Background
(06:26) Why IBM Set Up Research in India
(11:45) Can India Compete in AI
(15:18) How IBM Collaborates With Indian Universities
(19:25) Why India Has Been Slow in AI
(24:50) IBM's Hybrid Cloud AI Research Focus
(27:34) How Data Scarcity in India Makes Better AI
(31:18) Fine-Tuning Models Without Losing General Knowledge
(35:03) Continual Learning and Catastrophic Forgetting
(38:25) COBOL and Legacy Code Modernization
(42:11) Agentic AI Hype vs Enterprise Reality
(48:09) What Young Engineers Should Study Today




