In short
Notes on Podcast Episode: #232 Sepp Hochreiter: How LSTMs Power Modern AI Systems
Podcast Overview
- Podcast Title: Eye on A.I.
- Host: Craig S. Smith
- Episode Description: In this episode, Sepp Hochreiter, the inventor of Long Short-Term Memory (LSTM) networks, discusses the impact of LSTMs on AI, their architecture, and their applications in modern technologies.
---
Key Themes and Discussions
- Origins of LSTMs
- Background: Developed by Sepp Hochreiter and Jürgen Schmidhuber in 1991 to overcome limitations in recurrent neural networks (RNNs).
- Vanishing Gradient Problem: A significant challenge in training deep networks, where gradients become too small to effect learning as they propagate back through time or layers.
- Memory Structure: Introduction of memory cells that maintain contributions from earlier time steps in a sequence.
- Architectural Features of LSTMs
- Memory Cells: Ability to store and retain information over long sequences.
- Functionality: Designed to handle sequence data like text and speech effectively.
- Role in Technology: Pivotal for applications in major tech companies like Amazon, Apple, and Google; foundational for early large language models.
- Comparison with Transformers
- Transformation in AI: The emergence of transformer architecture in 2017 revolutionized the field, facilitating better scaling and parallelization compared to LSTMs.
- Limitations of LSTMs: Initially outperformed in language tasks due to the inability to handle large datasets as efficiently as transformers.
- Introduction of XLSTMs
- Development of XLSTMs: A newer iteration of LSTMs integrating Hopfield networks for enhanced memory capacity and performance.
- Goals and Advantages:
- Energy-efficient and faster processing.
- Scalable for larger models.
- Suitable for real-time applications, especially in robotics and industrial automation.
- Applications of XLSTMs
- Real-Time Robotics: Capable of controlling autonomous systems efficiently.
- Industrial Use Cases: Highlighted examples include controlling drones and time series predictions relevant to financial markets.
- AI for Simulation: Potential for numerical simulations to enhance speed and efficiency in various industries.
- Future Challenges and Opportunities
- Data Limitations: Discussion on the saturation of available language data impacting future scaling of models.
- Continuous Learning: Consideration of whether models should adapt post-training based on environmental changes.
- Exploration of Different Modalities: XLSTMs are proposed to handle different types of data, moving beyond traditional language processing.
- Community Engagement and Open Source
- Open Source Model: The XLSTM model and its components are made available for public use.
- Encouragement for Exploration: Hochreiter invites the community to experiment with XLSTM technology, emphasizing its potential to innovate in the field.
---
Conclusion
- Vision for AI: Sepp Hochreiter shares optimism about the future of AI technologies, particularly emphasizing the role of efficient models like XLSTMs in transforming various industries.
- Impact on Research and Industry: Discussion of the balance between theoretical advancements and practical applications in real-world scenarios.
---
Key Takeaways
- LSTMs are foundational to many AI advancements and remain relevant in various applications despite the rise of transformers.
- XLSTMs offer potential improvements in efficiency, scalability, and applicability in real-time systems.
- Ongoing research and community engagement are crucial for further developments in AI technologies.
---
Additional Resources
- Website for NXAI: [nx-ai.com](http://nx-ai.com)
---
These notes encapsulate the discussion points and insights from the episode, providing a comprehensive overview of the contributions made by Sepp Hochreiter in the field of AI, particularly through the lens of LSTMs and their evolution into XLSTMs.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Now you have a task to solve. Perhaps was it a man or a woman crossing the street and man was at the beginning of the sentence. Now you go one time step back and look up, was it man or woman? And then you go the next time step, but now the contribution becomes smaller by a constant factor. Perhaps the first time you go back, it's 0.9 times the thing. Then it's again 0.9. 0.9 times 0.9 times 0.9 goes to zero. Then you say, if you encounter the word man, the contribution is zero. You cannot see it. You do not recognize the meaning if you go back in time, the contribution gets less and less and less.
0:42And that was the problem. Ideally, the contribution would be the same. No matter whether the word is at the beginning of the sequence or at the end of the sequence, it might contribute equally to the answer. And that was the idea. First I discovered the vanishing gradient. When then the idea came of the long short term memory. Here the idea was, can I build something, construct some architecture? If we go back in time, it's scaling with a factor of one. My name is Sepp Hochreiter. I'm leading the Institute for Machine Learning in Linz. I'm a university professor and also I founded a company called NXAI, which is dedicated to industrial AI.
1:24I'm very well known in this community because I invented LSTM, which stands for long short-term memory. And I did this in my diploma thesis, where my supervisor, Jürgen Schmidhuber, said, yes, these recurrent neural networks do not work, they cannot store something. I only can memorize the last words if you say, a man in a red shirt crosses the street. Perhaps the only memorized crosses the street. But if you say, the street is crossed by a man in a red shirt, then you know it's a man. In speech, you can permute the words and it's the same meaning. Therefore, memorizing is important. It's also important for time series and so on.
2:26At this time, it was 1991. A recurrent neural network couldn't memorize. And then I was digging into the problem. And what I discovered was a problem which is now known as a vanishing gradient. The vanishing gradient was also the reason why deep neural networks were not possible because what I discovered, the gradient which you need for doing learning vanishes if you go back in time and therefore you cannot see what happens at the beginning of a sequence. It's the same if you go through layers of a neural network. Also here's a gradient vanishes and you have the same problem. Therefore, neither recurrent neural network were possible nor deep networks were possible.
3:19And this was the problem of the vanishing gradient. Can you explain vanishing gradient a little more for listeners who are not mathematicians or statisticians? Yeah, it's a kind of credit assignment. Now we have a task to solve. perhaps in my previous example was it a man or a woman crossing the street and man was at the beginning of the sentence now you go one time step back and look up was it man or woman and then you go the next time step but now the contribution becomes smaller by a constant factor perhaps the first time you go back it's 0.9 times the thing, then it's again 0.9. 0.9 times 0.9 times 0.9 goes to zero.
4:16Then you say if you encounter the word man, the contribution is zero. You cannot see it. You do not recognize the man. Meaning, if you go back in time, the contribution gets less and less and less. And that was the problem. Ideally, the contribution would be the same no matter whether the word is the beginning of the sequence or at the end of the sequence, it might contribute equally to the answer. And that was the idea. First, I discovered the vanishing gradient. But then the idea came of the long, short-term memory. Here the idea was, can I build something, construct some architecture? If we go back in time, it's scaling with a factor of one.
5:04meaning it's equally important. If you go one step back, it's equally important. If you go another time step back, it's equally important. And this was the LSTM architecture. It's called the memory cell. The memory cell was a construction where you can go back over time and the credit assignment, how much something contributes to the final answer remains the same. And I built an architecture around this idea around the memory cell, which is called LSTM. And now it was possible to store information of hundreds or thousands of timestamps. It was not possible before. And then LSDMware was used in many different applications.
5:48It was used at some point in every cell phone, in every Android cell phone, in every Apple cell phone. It was used in Alexa. It was used by Do and Alibaba. everywhere were LSDMs for language, for speech and language. And the first large language models were LSDM models. Today we heard a talk, the Test of Time Award, and this was Ilya Sutzkefer, one of the founders of OpenAI, got it, where he used an LSDM for sequence-to-sequence learning, which is actually translating. LSDMs were everywhere. And Google was at this point known as an LSDM company. Yeah. Can I just ask on Vanishing Gradient? And I'll edit this, so don't worry, because I'm a journalist, not a machine learning guy.
6:46Is that related to Gradient Descent? Yes. So Gradient Descent, you're working step by step to find an optimal, right? Exactly. But it's gradient descent. It's the gradient. You want to compute the gradient to do gradient descent. But what turns out is that the gradient is super tiny. And if you have a tiny gradient, gradient descent doesn't work anymore because the direction cannot be determined. Because it's so a tiny number, your gradient is too small to do gradient descent. Therefore, gradient descent does not work anymore. Yeah, I see. And yeah, so LSTM was really the original working neural network, right?
7:36There was, especially for sequences, for speech, language, there was only LSTM. Everybody used LSTM, all big companies, Amazon, Microsoft, Apple, everybody used LSTMs. And every cell phone an LSTM was. Yeah. And what fascinated me about your talk is that you've continued working on LSDM. And in AI, there are these ideas that are the primary focus for a period of time. And then another strategy, another algorithm, another architecture comes along, and that becomes the focus. but it doesn't mean i mean that's what i i talked to isabel gill about it doesn't mean the support vector machines are useless they're there and and there's still things that can research it can be done on them the same with uh with uh well boltzman machines ultimately led to uh deep networks.
8:47But can you talk about continuing that research and how frustrating is it to have the community then turn their attention elsewhere? In 2017 there was a paper, it's called Attention is All You Need. The idea of attention was there, but attention was always used together with LSDMs. You have attention mechanism and LSDMs. And there was this paper and said we can do only attention looking back in time without an LSDM. And doing this, they could paralyze the whole thing much better. They can push much more data into the learning algorithm. They can use larger models and then this architecture with attention is called Transformer.
9:38The Transformer technology replaced LSDM. But LSDM were still used. For example, OpenAI had this Dota 5 reinforcement agent, and it was a big LSDM network. DeepMind has this StarCraft AlphaStar network. It was an LSDM network. And this year, in a nature paper, method come out to predict floodings or droughts and this is based on an LSTM network and it's done by Google. You can have a Google app and you can look up, is your home in danger of floodings? Do you have to take precautions or is everything okay? And even the US government and the Canadian government, the official method now is this LSTM networks to predict floodings or droughts.
10:41Therefore, the LSDM networks still worked, but where they were not used was in language, was for the large language models. And we saw LSDM is a very, very good technology. It excels everywhere, but not in language. But the main reason was the data. You could not push enough data. And for XLSDM, that's a revival of the LSDM in language, we looked at, can we do the same as with the transformer technology? Can we scale it up? Can we build a very big network? And what are some limitations of LSDM? And we identified three limitations. One was it could not revise storage decisions. If I store something, I cannot later say, oh no it was not correct I want to store something else.
11:37This we did with a technique which we call exponential gating. The second thing is we built a huge memory. LSDM had only a very tiny memory only one number was stored and we extended the memory and now the funny thing comes the memory is a Hopfield network perhaps you aware there was this thing which is called Nobel Prize and And John Hopfield got the Nobel Prize for his classical Hopfield networks. And we used, in two other lanes, his ex-LSDM, a classical Hopfield network, because it's very fast and very efficient. But equipped with LSDM things like input gate, output gate, we merged Hopfield network and LSDMs to get an ex-LSDM.
12:26And now it has much more memory. And the third thing was to make it parallelizable, like the transformer. With these three ingredients, exponential gating, having a good, efficient memory, and making it parallelizable, we saw we can be as good as transformer in, for example, the 7 billion model regime. But we have a big, big advantage. our advantage is we are much cheaper cheaper in terms of energy or in compute and we are faster because we don't have to compute so much we are also faster we are faster and more energy efficient both in training but especially in inference if you use it, if a user is using it we can accommodate more users at the same time or we can
13:24have one user but much less compute with much less energy. And this is a huge advantage of this new technology compared to the existing technology, which is a transformer. But what I also have to say for XLSDM, language is not our prime target. Our prime target is using this technology because it's so powerful to go into industrial applications because language is not at the core of many companies. It's nice for customer relations or PR or stuff like this. But with XLSM, which is so fast, we can go into robotics. Transformer cannot go into robotics. So I tried, but we're not fast enough. We have two advantages.
14:15We are fast enough to go in robotics, and we can fix the memory because we have a fixed memory compared in contrast to transformers. And with this fixed memory, we can adapt it to the hardware which is available, which is a smaller hardware on robots. And therefore, we can go to robotics, to embedded systems like drones, to embedded systems like autonomous production systems. and here we know it's the core of some industries. Here it's the core of some businesses of companies. And here we see a huge, huge advantage because we a priori know how large the memory is and we can go to the embedded system on the machine and we are fast enough to do it real time that we can really control a robot or autonomous thing in real time.
15:16Is it a light enough? In XLSDM model, could it be light enough to have on the edge? That's where we want to go. The guys from Apple ask us and so on. We have to look into it. What are the requirements? I think yes. If something can do it, then it should be XLSDM because it's memory efficient. And also the energy efficiency, if you have it on your cell phone or on the edge, you don't have these huge batteries with you. And if it's more energy efficient, it's also super helpful to use it. And therefore, I think that would be a main target for this kind of models. Because also look at the self-driving cars.
16:03You cannot have these big batteries. We are in contact with some car companies. that would be an ideal system for us to have it in the car because we don't have so much energy in the car because we have only a battery and we should work in real time. We should react very fast. Both we can accommodate. Both we can do. And I'm super optimistic that we have a super cool thing here. Yeah. Going back to attention. So attention existed, and forgive me because I don't know this history. Where did attention come from? Who developed that? So Transformer was 2007. Attention is all you need. But the attention idea, it was always in combination with LSDMs.
17:00There were different developments. They were something like, let's have an external memory where you access the memory. Also, Alex Grave had an DSM, which was called the neural Turing machine, an LSDM, which is a Turing machine. And from this direction, there was this hard attention. I store something and I really get an address to what I should attend to. this was Jason West memory networks the idea was memory networks it was always can I store some information LSDM has something in its memory but at this time the memory was very small and they wanted to have a bigger memory. Now with XLSDM we use the Hopfield network but they used a separate memory mechanism in addition to LSDM to store more information from the past.
18:00And there were different ideas out. And then there were addition attentions. There were multiplicative attentions. There were different attention mechanisms. It was always, can we integrate into LSDM a more powerful memory? But they didn't replace a memory cell. I started with a memory cell with this constant not-vanishing gradient, constant gradient going back. they didn't touch it because they don't want to destroy that. But you can build a memory directly into LSDM. This is what we did. The others did a separate memory. Their attention came. And since I saw the separate memory is powerful enough and we don't need the LSDM anymore.
18:46This was 2017. Attention is all you need. But in 2015, the first attention ideas were already out. I see. And so the XLSTM, where you have an LSTM and a hot field network to expand the memory and it's faster, how is it trained when you say a 7 billion parameter model? Is it trained in the same way that a large language model is trained? It's exactly the same. We use the same backbone. And as our attention mechanism was located, we now replaced this with an LSDM. But with an LSDM, it's a huge memory, exponential gating, and with the possibility to parallelize it. And we did the same. We looked what the transformer guys are doing, how do they train their networks.
19:44We did the same. We have to do some adjustments because we have a specific architecture, but the major thing came from the attention backbone, which is a ResNet. We only replaced in the big network this one item, which is quadratic in context length, also is linear, and we are much faster by replacing this compute-intensive part by a less compute-intensive part. When did you do this work? Is this very recent? It's very recent. And as attention came up, as I saw it, I thought, oh, probably we can do the same with LSDMs. But the development with transformer attention was so fast that we never had the chance to do something.
20:39And then the problem was the models became so large that we at the university did not have the compute. Then I looked at some developments and sat down, computed everything through with this Hopfield Networks, with this idea. And then I went to the press and said, I have a new idea, XLSDM. I need money to try it out. And then after a while, there was a local or private money, I think, which gave us money. and we founded the company, the NXEI company was founded to build the XLSTM. And the first month we invested 10 million euros only in compute to see whether the XLSTM works. I cannot do this at the university.
21:32I needed private money. And it worked very well. Then we said, let's build a larger model. Let's go to the 7 billion model. And it worked better than I dreamed. It's so fantastic. and I'm so proud I had the first LSTM where all the big companies have used it but I never saw any send nothing now I have the second chance with ex-LSTM but in the company we have the IP rights and stuff like this we try to use it on our own and bring it to industry but it is so fantastic that it's really working like a dream Yeah, that's amazing. And you say it's not particularly, it's not built for language, but can it handle language?
22:23Yes, it's as good as this language models, like, you know, LAMA, LAMA 2, LAMA 3, it's better. Yeah, but it's also much faster. It's more energy efficient. But in the meantime, there are better models than Lama. It's in the middle field of these language models. But language, there are so many models out, so many things. This is a big chance to go in industrialization, to have a very fast, real-time method, energy efficient, to really go to the core business. Because with language, you're often not at the core business. You don't make the big money. Yeah. Yeah, so give me a use case, an industrial use case for XLSTM.
23:17What was already done, not by me, by others, some implemented XLSTM in controlling drones. And it worked fantastic. As I said, it worked better than the transformer. It worked better than the state-based model. it worked for them very good. It was not my idea, so I tried it out. Others used it for time series prediction. This was Wessler or even financial time series. And they reported in a paper, they published it, that they're much better than Transformer, much better than classical time series method. and this was another I'm very proud another reinforcement of that this method really works well and on archive a couple of papers as I tried it out and they always get new state of the art new good results and now I hope it's getting more and more popular more people are using it and see the power of this new method I'm so excited Yeah, I can imagine.
24:33Particularly since you had this huge success and then Transformers kind of eclipsed it and now you're coming back. Is it open source or how are you approaching that? we made everything open source meaning the model is open source the inference code is open source that you can use the model but also the training code is open source everything is open source and now the funniest thing comes we are very proud we are faster than flash attention I don't know whether you know flash attention it's a very fast transformer technique and it turned out that we are faster and being faster is because we have this recurrent neural network, we have an LSDM what we did is we have chunks of flash attention and if we can decide how large the chunks are you can better make use of the GPU of the graphic processing unit of the chip if you have to do everything in one big flash attention thing you have to squeeze it in, you cannot do the optimizing but we have more smaller chunks and can optimize the chunks for the GPU.
25:50And now we are faster than flash attention because everybody thought, I told before, a transformer become popular because they could be paralyzed and they are faster. Hey, and now we are faster. We are faster than the fastest transformer. It's unbelievable. I never thought this is possible. But we achieved it. We really achieved it. On the drone, excuse me, on the drone example, is this model on the drone or is it communicating with the drone remotely? And what exactly is the model doing? Is it guiding the drone? Is it guiding it through obstacles? Can you talk about that application? I'm not a part of the team.
26:47This is somebody else. I was only told as far as I know, but I can be wrong. It's on the drone, on the device, on the drone, and it's controlling the drone without external remote. They have probably remote for safety reasons, but the idea is the whole system is on the drone, that you have an autonomous system. But it's a different lab. They did not publish this until now. They only told me that we were very successful, and that's the only information I have. Yeah, and so NXAI, is that how you're pronouncing it, or do you call it NXAI? Yeah, we are not sure. Until now, we call it NXAI. Elon Musk has something which is called XAI.
27:38Perhaps we should call it Next XEI or New XEI. I don't know. But right now it's NXEI. That's a company dedicated to industrialization of AI. It started with this XLSTM. We needed money to show such a cool technique. And as I told, it really worked out more than I ever dreamed. But we have a second pillar. We have a second thing, and that's AI for simulation. It's also in the company. AI for simulation is you have these numerical simulations. You have many particles or you have a mesh over your car where the air is flowing. and these numerical simulations are restricted. If the number of particles goes into millions or hundreds of millions, so numerics become too complicated.
28:40You have to do too many computations. But AI can be faster, can be 1 ,000 times faster, can be 100 ,000 times faster because AI can see meta structures. I'll give you an example. You know, the moon is circling the Earth. The moon has many, many particles and so on. But we describe the moon by its location, by its impulse, perhaps its mass, only a few parameters. And we have a very good prediction where the moon is tomorrow in an hour. And this is a very high abstraction. And many physical systems have, let's say, a set of particles which form one structure, like quills or other structures. And AI recognizes the structure.
29:38Well, I only think you throw a snowball. You don't have to simulate every snowflake, but the ball. And if the AI recognizes, hey, this is a snowball, let's simulate the whole ball as one, then you're much faster because 100 ,000 particles, 1 million particles are now one thing. And the idea in simulation is AI recognizes macro structures, simulates these macro structures, and it's much faster, but has almost the same precision as a numerical simulation, which goes down to perhaps atoms. and this AI for simulation is super important you can do it for the food industry if you store something like in a hopper in some compartments how it's flowing out you can do it if you have some particles and you blow air into it to mix different kind of particles you can see how the mixing makes good even if the particles have different sizes, different structure, and all kinds of things.
30:52You need the simulations. In Linz, we have a big steel industry. For this big office of the steel industry, the simulations, the numerical simulations are not sufficient because there are so many particles, there are so many atoms to simulate that it doesn't work anymore. But AI can do. And here we see with the simulation, a huge field in industry, at the companies where we can do simulations, where numerical simulations are structured, where numerical simulations are not fast enough anymore. That's the second thing. XLSTM, it's a new technology. It can be used for everything. And simulation using AI.
31:40This is the two main pillars we have in the new company, and XCI dedicated to industrial AI. You know, I had a conversation with Ilya Sutskiver on the podcast, and he was saying that the transformer scaled very well, but it could have been anything that scaled well. He was saying that they worked with the transformer because it did what they needed to do, but he felt that it could have been a different core algorithm. Is that what you're finding with LSTM? Exactly. At the beginning, it's the same feeling. This is a tension at the core on LSTM or something else. it's the scaling up building a very huge network storing all the information in the weights, in the connections I think that's the important thing to build up a big network they use transformer it's only disadvantage now everything else would perhaps also have worked it's quadratic in the context links they used a little bit more complex I think and it's complex in time but it's simple in the interaction because you have a query and you have keys.
33:09Every query looks at one key. You have only pairwise interactions. With LSTM, you have everything in a memory and a new query interacts with the whole memory. Certainly, you can have also more complex, more abstract interactions. But I decided for this, I think others also go. But now if you look in detail, perhaps it's better to have something more compute efficient something which does not need so many resources in terms of computational cycles and so you've scaled this so far to 7 billion exactly but there's no reason if it's faster and more compute efficient presumably could you scale it beyond what they've done with LLMs.
34:05Yes, I think so, that we can do it. We have to see whether we have enough money, enough compute to do that, because we are looking for investors and so on. We need money to do that. But I think we're in a good place, because we are now faster. We can do even larger models. But the question is, how large should the model be? also earlier today said perhaps it has an answer scaling. Also in my talk said there was a scaling. There's no more data available. We used all the language data up and producing with existing methods, new language data also doesn't work because it's not new information. It's old information.
34:49All language only transformed. So we don't have new data. and here I think we are saturating because we used up all language data we have also from Chinese, Japanese, across the world, this was all translated. If we want to build larger models, but we need also more data. But the data is a limitation. We don't have more data. And I think this is an answer scaling. Now we have to think more in inference and OpenAI has this strawberry model where you should think more or what I think it's more there's information in the model but you have to get access to it you have to tweak the model a little bit to get out what you want to have and that's an inference and now the cool thing is XLSM is fast in inference exactly here we have an advantage you can do much more if the transformer can do 100 tokens, we can do 10 ,000 tokens at the same time.
35:58And here for inference also we would have an advantage. And therefore I'm very hopeful perhaps still being in language. I said we want to go to industrial application but language is interesting after the 01, after the strawberry of OpenAI. I'm super super
36:17hopeful that we have a system which can think much faster. And as a system of direction, I want to look into it. Do you think that XLSTM has promised to challenge the large language model paradigm, where in a few years we'll be talking about XLSTM models instead of large language models? Depends. I don't know how everything has worked because this existing technology is very good integrated in the hardware. Nobody wants to change a working system. but if I say you save a lot of money because you need less compute I don't know whether we are too late to change because everybody bought already their language models had their product on the market I don't know but I think there are certain fields we have an advantage but whether we can change things which are already fixed I don't know But in principle, it would be possible.
Read the full transcript
37:36You have a model. It's a GPT. It's a transformer model. Now you replace it by an XLSTM. And you have the same performance, but pay less energy costs. But I don't know. It's a market thing. And I'm sure. On the training day, though, you know, I heard Yann LeCun speak yesterday. and he was talking about this problem and how we need to develop methods to collect data from the real world, not filtered through language, through various sensors, through real-time video, through haptic sensors, through, I don't know what else there is, but audio, I suppose. Could XLSTM handle different modalities? Yes, absolutely it can.
38:40And that's also with simulations where we go. With simulations, we also need data. There comes data from numerical simulations, which simulates nature very precisely. Then we can generate more and more data very slowly because AI methods are faster, but if you wait enough, then you have enough data to build them. But also from sensors. If you want to use XLSDM in different applications, we rely on this sensor data. Exactly like Jan said, I think that's the future. That's where we get new data. We don't get new data in language, but we get new data in many different applications because we can build more and more sensors which just record what's happening in the world.
39:33For a machine, how a machine is in the world. And then we can do the same as we do with languages, use like an Excel SDM to build a good model. And now a model which can control a machine, can make decisions or stuff like this. I think that's the future, yeah? And the learning, are these models, do they continuously learn after the training phase or are they fixed? That's your decision. You can do both. But normally what you do, you fix them. Because if the user is training these models, you don't know what's happening. It could be very risky. if you can also train a car to stop at green and drive at red because if you're a very bad teacher to the car and for some safety reasons, you don't want them to learn.
40:36But on the other hand, there could be machines where the machine dynamics change over time because some parts are worn out or not precise anymore. But then the model can adjust to the machine or adjust to the weather. Perhaps it's summer and then it becomes winter. It could have advantages, but one has to look in the application. You can let it learn, but you have to see it's dangerous that the user is in control in learning something. Yeah, well, I wasn't thinking about consumer-facing applications, but for example, industrial robots. Yes. You know, again, Covariant, this company founded by Peter Abiel, their dream is to have, and I think they already have them, networked.
41:35these industrial robots deployed around the world and as they encounter new situations and learn to deal with new situations, that learning then is sent back to a central model and then deployed to all of the robots. So you know this learning. Is that something that could happen with XTLSM? Yes, Excel STEM can do that, but it can do it even better. Right now, with robotics, you cannot learn on the robot because these systems are too big to climb them. But with Excel STEM, you can do it embedded. You can build it on the robot, and an individual robot can even learn. You can use the data, pull it, and learn a new big system.
42:27But with XMSM, it would be possible that each robot learns, adapts to its own environment, to its own task. And that's the thing. What we have in mind, for example, is if you have the industrial robots, you have over 100 ,000 millions of items to build. but for many smaller companies you have only perhaps 500 or 700 now it's it's not you don't program a robot because it doesn't pay off to program a robot if you have a robot which can learn perhaps you show the robot you have to do this screw in here do this do this you show it on two or three examples, and the robot does the other 700 examples.
43:23Here, XLS 10 would be good. It's on the robot. It's a foundation model. It knows how to screen and do stuff, but now you have to say for this item, you have to do exactly this. There will 700 of them, and it's doing this. That's not possible right now because an industrial robot can do it, but the program costs more than doing what a human is doing. And here, with XLSTM, we can put it on the robot, and therefore, it can learn more in context, because you would do the demonstrations from human into the context. And now, with this in context, it will do a smaller number of pieces. And here, I see also a huge potential.
44:16Yeah. Is this deployed already on industrial robots? No, no, because we come out with a language model. We first wanted to show our new technology is powerful enough. Our new technology is as powerful as the existing technology, which is called Transformer. Now, if you have shown this and everybody believes us, we will go into these different applications. And NXAI, when was it founded? It was founded end of last year, December last year. Yeah. And how much money have you raised so far? We have internally, I cannot reveal everything, but a couple of tens of million euros. but now we wanted to start a second founding round.
45:16We didn't start a round to get new investors in. We now first put out the 7B model and seeing how it's going and now we hope money comes in to go the next step. Because now we have proven we can do it and I'm really, really optimistic that this really goes through the roof. Yeah. Are you satisfied with the attention you're getting from the AI community on this LSTM? I mean, I was really impressed by the talk that you gave, but I have no way of judging. Yeah. Perhaps you have a better way than me. I don't know because also I'm in the bubble. Everybody who's coming to me says, hey, a cool thing.
46:11And also I used Excel as damn, it's really good. But I don't see the big picture. And we have to see how many downloads we have, how many publications come out, how many other success stories are there. If there are more and more and more success stories, it worked. But we have to see. I'm not sure whether it really grows or whether it's, yeah, I don't know. Sure. Do you think that there are other algorithms out there yet to be discovered that can scale and are maybe even more efficient? I think the field is open again. There's also the state-based models like Mamba. we compared to Mamba, we have advantages to Mamba.
47:02And the nice thing is Mamba 2, the newest version of Mamba, is an XLSDM without the exponential gating. We converge to the same architecture even. We have a little bit more performant architecture. But this was also funny because they come from this direction, we come from another direction. But we discovered the same things which work. but there could be other things out there and also a building on top of XLSDM. It's quite new. A handful of researchers have developed it. We don't know what is the potential, in what direction it can go. Everything is possible here. Mr. Transformate was the same.
47:47It was discovered, published and then many, many, many ideas from the researchers come in. And here I think also there's a huge potential. For example, you can modify some memory. For the Transformers, there was something like, it's called prompt engineering. Here you can modify some memory. If you say, no, you were friendly in the past. No, you were very technical in the past. You can directly modify memory. I don't know what's in there. I know it for biological application, what they did. So I learned certain proteins. put it in the memory, and took the memory to generate something else. I can take this memory, this memory, or a combination of memories.
48:32I scan over, store information in my memory, and then I pull it out. There are complete new ideas possible here. Is there anything I haven't asked that you think I should ask? Oh, I don't know what the audience is interested in. The main message came over. I hope that there's a new technique. Everybody can use it because it's open source. Everybody should try it out, play around with it, and you will have a very rewarding time to check it. I don't know what's missing. Yeah, yeah. And on the reasoning, you mentioned Strawberry or O1.
49:26Will you develop a reasoning capability, this sort of long-thinking capability that OpenAI is working on? Yes, we also think growing in this direction. I did not reveal what they're really doing, but they also from deep mindful other research groups some ideas out how to do the reasoning how to do the thinking we will also try out different things which are out there because we see an advantage that we are so much faster in inference and this should help us but we will try out many things and if OPI once revealed what they have done perhaps we also build it Yeah. On your site, is there a public-facing demo?
50:21Not now, I think. It's because we built a 7B model, but did not have time to find you. You need to find you to a chatbot or whatever. That's not what we had time right now. We are super proud that we got this cool model. Yeah. Okay. So we've got to watch you develop. Yeah, please do it. NXAI, what's your website? NXAI. It's, oh, I don't know. I don't have some mind. NXAI? No, NX-AI.com. Yeah, NX-AI.com. Yes. Yeah, well, I'll follow it. Okay, this is fantastic. I'm really glad that I got you up here.
From the publisher
In this special episode of the Eye on AI podcast, Sepp Hochreiter, the inventor of Long Short-Term Memory (LSTM) networks, joins Craig Smith to discuss the profound impact of LSTMs on artificial intelligence, from language models to real-time robotics.
Sepp reflects on the early days of LSTM development, sharing insights into his collaboration with Jürgen Schmidhuber and the challenges they faced in gaining recognition for their groundbreaking work.
He explains how LSTMs became the foundation for technologies used by giants like Amazon, Apple, and Google, and how they paved the way for modern advancements like transformers. Topics include:
- The origin story of LSTMs and their unique architecture.
- Why LSTMs were crucial for sequence data like speech and text.
- The rise of transformers and how they compare to LSTMs.
- Real-time robotics: using LSTMs to build energy-efficient, autonomous systems.
The next big challenges for AI and robotics in the era of generative AI. Sepp also shares his optimistic vision for the future of AI, emphasizing the importance of efficient, scalable models and their potential to revolutionize industries from healthcare to autonomous vehicles.
Don't miss this deep dive into the history and future of AI, featuring one of its most influential pioneers.
(00:00) Introduction: Meet Sepp Hochreiter
(01:10) The Origins of LSTMs
(02:26) Understanding the Vanishing Gradient Problem
(05:12) Memory Cells and LSTM Architecture
(06:35) Early Applications of LSTMs in Technology
(09:38) How Transformers Differ from LSTMs
(13:38) Exploring XLSTM for Industrial Applications
(15:17) AI for Robotics and Real-Time Systems
(18:55) Expanding LSTM Memory with Hopfield Networks
(21:18) The Road to XLSTM Development
(23:17) Industrial Use Cases of XLSTM
(27:49) AI in Simulation: A New Frontier
(32:26) The Future of LSTMs and Scalability
(35:48) Inference Efficiency and Potential Applications
(39:53) Continuous Learning and Adaptability in AI
(42:59) Training Robots with XLSTM Technology
(44:47) NXAI: Advancing AI in Industry




