RWKV: Reinventing RNNs for the Transformer Era — with Eugene Cheah of UIlicious

30 Aug 2023 · 1 h 12 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Latent Space Podcast Episode Notes

Podcast Information

  • Title: RWKV: Reinventing RNNs for the Transformer Era
  • Host: Not specified in the transcript.
  • Guest: Eugene Chia, CTO of UIlicious and core member of RWKV committee.
  • Date: Not provided in the transcript.

Episode Summary This episode delves into the RWKV model (Receptance Weighted Key Value), which aims to challenge the dominance of Transformer architectures in the landscape of large language models (LLMs). The discussion highlights the potential of RWKV to address critical challenges in LLMs, such as improving context handling, reducing computational costs, and introducing a novel architecture that revives recurrent neural networks (RNNs).

Key Concepts and Discussions

  1. Background on RWKV
  2. Introduction to RWKV:
  3. RWKV is a new architecture that seeks to combine the benefits of RNNs and Transformers.
  4. The model has been trained with up to 14 billion parameters and aims to provide competitive performance on reasoning tasks while ensuring linear scaling in computational costs.
  1. Challenges with Transformers
  2. Context Length Limitations:
  3. Traditional Transformers scale quadratically with context size, making them inefficient for larger contexts.
  4. RWKV claims to alleviate this issue with linear scaling characteristics.
  1. Community and Development
  2. Open Source and Community Focus:
  3. The RWKV project is primarily driven by a distributed, volunteer community, similar to Eleuther AI.
  4. The community is polyglot and seeks to support multiple languages, driven by user needs rather than benchmark performance.
  1. Technical Insights
  2. Architectural Innovations:
  3. The RWKV architecture replaces traditional attention mechanisms with time and channel mix operations, allowing for efficient parallel processing and memory handling.
  4. Scaling and Training Efficiency:
  5. Discussed the memory efficiency of RWKV and its design to handle large context sizes better than current Transformer models.
  1. Current Limitations and Future Directions
  2. Model Limitations:
  3. Despite its strengths, RWKV struggles with long-distance dependencies and may require advancements to fully exploit larger context sizes effectively.
  4. Future Plans:
  5. Eugene highlighted ongoing efforts to refine RWKV's ability to handle increasingly large contexts while maintaining efficiency.
  1. Career Advice for AI Engineers
  2. Paths for Engagement:
  3. Emphasis on the importance of being mercenary with tools and models, focusing on practical application rather than theoretical complexities.
  4. Encouragement for engineers to experiment with model training and fine-tuning, and to leverage existing resources and communities for learning.

Timestamps of Discussion Highlights

  • [00:05:35] Eugene's path into AI at UIlicious
  • [00:10:17] Limitations of Transformers for large context sizes
  • [00:24:45] Overview of RWKV and its architecture
  • [00:53:29] RWKV's competitive performance on reasoning tasks
  • [01:03:10] Advice for AI engineers wanting to gain technical knowledge

Key Takeaways

  • RWKV presents a significant alternative to Transformers by reviving RNNs while addressing their limitations.
  • The community-driven nature of RWKV fosters a rich environment for innovation and multilingual support.
  • The episode emphasizes the importance of practical engagement in AI and learning through experimentation and community interaction.

Additional Resources

  • RWKV Documentation: [RWKV Docs](https://wiki.rwkv.com/)
  • RWKV Model GitHub: [RWKV on GitHub](https://github.com/BlinkDL/RWKV-LM)
  • EleutherAI: [EleutherAI](https://www.eleuther.ai/)
  • Recent Papers and Models:
  • [Attention Free Transformers paper](https://arxiv.org/abs/2105.14103)
  • [Retentive Network paper](https://arxiv.org/pdf/2307.08621.pdf)

Conclusion The episode provides a comprehensive overview of the RWKV model, exploring its innovative approaches to address longstanding issues in LLM architectures while highlighting the collaborative efforts of the AI community in pushing the boundaries of the technology.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hey, listeners. Today, we have a very special episode for you. There's been a recent paper on the top 10 open challenges in LLM research that has consolidated a lot of intense debate. Today, we're going to talk about RWKV models, Receptance Weighted Key Value Models, with Eugene Chia, who is both part of the core RWKV team, CTO of a low-code AI test automation platform, and active member of the Latent Space Discord. The RWKV architecture has the potential to solve three of the top 10 open LLM challenges, increasing context length, making LLMs faster and cheaper, and designing a new model architecture.

0:32What is particularly appealing about it is that it does so by reviving the recurrent neural network, which even I have argued has been obsoleted by the transformer. It rejects the idea that attention is all you need, and replaces multi-head attention and feed-forward networks with new concepts called a time mix and channel mix, respectively. It has been trained up to 14 billion parameters. They're getting help from Eleuther and Stability AI to scale it up even more. It shows competitive results on reasoning benchmarks, the same benchmarks we covered in our Benchmarks 101 episode, with similar size models, yet with linear costs and speed curves instead of quadratic ones.

1:04In a way, RWKV are promising the room temperature superconductor of LLM architectures. In other words, the parallelizability and performance of transformers without the quadratic cost. Obviously, the topic of what happens after the transformer is in the finance terminology, what we call a low delta outcome. It's probably not going to happen, but if it does, it will be very, very big. We even discussed this a little bit in our episode with Jonathan Frankel of Mosaic ML. But since the RWKV paper was published, the idea has been somewhat independently validated with Microsoft Research putting out the RETNET, or the Retentive Network, which has similar veins of what it's trying to do.

1:41And it also, of course, competes with other alternatives to the transformer, like the state space models coming out of Chris Reyes' group at Stanford, S4, H3, and the Monarch Mixer that was recently announced. However, RWKV is so far the most validated of all these ideas because it is already trained up to 14 billion parameters with multiple models that you can download and generate text with today. As podcasters, we want to be the first place that you hear about new things in AI that you'll be using in work or personal life as AI engineers and enjoyers. So this presents us with a problem. We have to be early on consequential topics and things, but also high signal.

2:15Some of our favorite compliments so far, which by the way, I've added to our about page if you want to check that. Your pods are a legitimate highlight of life for me. They're amazing. From McKay Wrigley and from the AI Safety Memes Twitter account, which is always fun. They just simply said we're the highest signal pod for them. So we're very proud of this and want to keep it up while taking risks because even though we crossed a quarter million downloads just five months into the podcast, we're still very young and trying to figure out what kind of podcast we want to be and what kind of audience we want to have.

2:44Today is going to be one of the riskier pods for a few reasons. One, it is the first pod we're doing on a non-traditional architecture with no large western institutional backing. Two, it is the first pod we recorded outside of the US without a regular studio, so the audio is not as good. And three, finally, it is the first pod we're doing with my Singaporean accent. I'll address each of these in turn. One, there's a significant institutional bias in western media coverage when it comes to AI in that if you don't come from Stanford or Oxford or you don't work at Google of Facebook, your ideas have trouble getting attention.

3:18However, by far, the most admired organization that our guests have repeatedly mentioned is EleutherAI, which spun up as a decentralized Discord community that independently trained the first GPT-class LLMs without a prestigious background. Eleuther has since spun off organizations like StabilityAI and Conjecture, but I suspect that the RWKV community is working like the early days of Eleuther, and we have a rare opportunity to capture an oral history of it live rather than two years after the fact. Two, part of the latent space magic is that we try to get to know our guests and be in person with them to establish rapport like we did in New York in our Notion AI episode with Linus Lee.

3:53Despite the lower audio quality, I think we got a much better interview with Eugene because we were able to interact with each other in person. Three, part of the joy of audio is that you get to hear the full diversity of humanity and neither Alessio or I are American and I like to showcase the broader world through our AI lens whenever we can. As you'll see, the RWKV story is also about the non-English rest of the world, self-organizing to build the LLMs that the English-centric West has neglected, and doing so relatively successfully. We've aggressively edited this interview down to one hour for the audio podcast, but for those who are interested in RWKV, head over to the new Latent Space TV YouTube channel, which has the full two-hour interview, including our screen share while through the RWKV models paper and Discord, as well as digressions on Eugene's background, discussions on diffusion models and the token crisis and the relationship between open source ai and the ai waifu hasbando community i hope you enjoy our conversation with eugene chair of rwkv and if you liked it reach out to him at pico creator on discord or on twitter and let him know okay so i'm here with eugene we are in singapore we this is the first time i'm podcasting in singapore this is the first time i'm podcasting with my singaporean accent eugene has been a very valued part of our latent space discord for a while and also diving deep onto rwkv i think you're actually the first person that brought it to my attention as like a potential transformers alternative you're also cto of uilicious which is a ui testing company that's in singapore here which is local platform i got the first demo maybe four years ago yes and i was like okay fine you know you're doing testing there wasn't an obvious ai angle i mean now that you explained it it was great but like what was your personal like Okay, I'm gonna be a dedicated AI guy for UiLashes.

5:38Okay, so one of the things that I found very interesting with the huge Transformers boom right now is that traditionally, right, when you tell companies that you want to build your own AI, you need a really large data set. And over time, actually, the amount of data sets that you need is actually scaled down because you can just now find... Foundation model. Yeah, 5.1 Foundation Models. And when we started Neuralicious, we always knew at that time, because a lot of our other companies that were launched at the same time were dealing with Neural Networks, that at some point, the data that we've been collecting data on, let's say, how to do testing, website, it's just a very specific focus.

6:15Basically, every single test that has run on our platform, unless our customer has opt-out or deleted their account, basically privacy-related stuff, we actually still retain the test data. and that's something that we always felt that was useful in the long run to be able to actually build a huge training model. The irony of that was that even though we were building all those data sets, as the threshold came in and the transformer boom happened, we realized we don't actually need that big of a data set anymore to actually get a functional AI. One of the key insights, especially for people who is like trying to build on top of a transformer model, pre-transformer large language model is we would always be thinking of like in terms of like 100 gigabytes of data, one thing to buy a data, millions of records for all the different examples.

6:59Post-Transformer, it's literally, you probably need only like a thousand or ten thousand, enough data that you can literally get an intern a few weeks to just get it done. And you have a working model. It may not be that great, but frankly, every piece of data you add after that is a diminishing return. Because it's a language model, it doesn't actually have any inherent understanding that it is automating the browser. So it's presented as like a prompt answer pair, like question answer pair. At least for our internal model that our users are using, it's presented as here's the prompt, describe your task, or what you want to modify the code, and then subsequently generate the code for you.

7:37And hindsight is now basically copilot. I think now copilot is adding that chat. Rigid, are they fully launched yet? Yes, I actually downloaded it yesterday. I haven't actually used it yet, but it is a separate VS Code extension. So there are now three co-pilot extensions shipped by GitHub because they have shipped their org chart. I'm quite friendly with that team, but it's very funny. But just to come back to you, so did you implement this with GPT-3? So we based it off the Salesforce CodeGen model. Okay, right. So that was the foundation model that we built on top. We are looking into replacing it in parts, but that becomes a longer conversation.

8:14CodeGen being the first really credible open source code specific language model that was released by literally anyone, I think about three years ago. Yeah. And then they recently released CodeGen 2. Correct. Any opinions on CodeGen 2 while we're on this topic? In terms of CodeGen, one big appeal for the CodeGen and even CodeGen 2 model is that Salesforce took a very clear and clean approach to the licensing. Meaning they were very, very clear that everything that they trained on was open source. Yeah. MIT, they didn't touch the problematic languages. So, and you can imagine... And you think that Copilot did?

8:53Knowing Microsoft's statement on how liberal they were about GitHub data, and they were saying they used the term that it's under fair use. I see. Yeah. I have no reason to believe that they didn't. but this same problem happens to actually a lot of existing code jam models and and that was actually the main appeal for me for running for actually building on top of the salesforce code jam model mostly also because like like for us we deploy on-premise into enterprises yeah in europe and they ask questions so what what does this deploy on-premise mean like you you pack ui lishes into a container yeah you give it to them yeah and then it's like a license fee or something Correct.

9:35Okay, cool. That's very interesting. Yeah, okay. I don't know if I have any other questions based on that. Anything else before we go into the reasons for alternative models? Okay, so anything else do I have there? No, I don't really have much, right? For alternative models, so yeah. So let me just set the premise, right? Transformers have won, for now. They have slid the neural networks. Yes, and it seems like you have had a history since with machine learning, since before Transformers, and now they're at kind of the peak of their power. And I see that, you know, there's a desire for alternative models for a number of reasons, but I'm very curious as to what drives your personal interest in alternative models.

10:18So first things first, to be clear, majority of AI is still based on Tranho, at least within my company. But what drove me into alternatives beyond Transformer, in essence, once we actually managed to get our bot to generate UI testing code, The most obvious next thing that our customers started asking, hey, let's say the test failed. Can your AI now analyze my website and then tell me what's wrong and tell me what to change? Basically, they're getting lazy and lazy. Yeah, yeah, yeah. Humans are very good at moving GoPo's. And we had something working for toy websites. But the one thing that we do internally is that we look at the, I think, what was the list?

10:58Top 100, top 1 ,000 websites. And we basically just run, or we actually do run our test platform against that to see, make sure that our code works against any front-end platform. Well, what do you mean run your test platform, right? Because you don't have tests for them. Yeah, we have some very rudimentary basic things. Like, go to a website, see something, click something, add to cart. Yeah, that's it. The idea is more of like, because there's so many frameworks out there. You just want to make sure you cover all of them. Yeah. And so we did the same thing for our AI. And the first thing that it died on was literally Amazon.

11:31Why? Oh, 5 megabytes. Yeah, I think you heard me mention that. So, when you are trying to analyze a website, we've been talking about increasing token count size, right? But for e-commerce websites in particular, even if you strip off a CSS, even if you strip off a JavaScript, having the entire HTML in megabyte size is not unheard of. Yep. And that's where it's like, how am I supposed to solve this in terms of an AI point of view? How many tokens would that be? Like, oh my gosh, you could easily be looking at over a million tokens. I see. Which is still too much even for today. Yeah. Did you look into making your own tokenizer?

12:08That's something that we explored. I think what we found more realistic was to actually pass the HTML into a more token-friendly format. Yeah, right. So this way we can still build on top of existing models. But yeah, we are exploring that as well. But back to the alternative. So the key thing for me was at that point, and subsequently I think I showed you the experience with English compiler and things like that, AI agents generating code, you also have your own small there, was that the context size is a real problem and Transformer inherently by its nature, at least vanilla Transformer, I know that's Transformer XL and some other attempts, is that it quadratically scales with the contact size.

12:58So if we scale to, let's say, 100 ,000, that's already requiring a shit ton of compute and re-ramp. And I don't even want to imagine what happens to 1 million or 10 million. And that's where I needed, I was like, okay, this is a fundamental problem that needs to be changed. If not, we will not go past this. and I think there's also now a lot of people who are very interested in models that can handle large context size because they also want it to be able to use in use cases where they never need to fine-tune fine-tuning is a pain, apparently. Yes, that said there's issues with just throwing everything in context, right?

13:38It's shown that retrieval is only best when the item that's relevant is in front or in the back of the context window so it's basically i'm just like maybe we've just tapped out context is working memory and maybe it's like maybe transformers are very similar to humans in that a working memory is only of a given size if you try to artificially extend it you just have you just make it very lossy yeah so so so that's where i end up landing on the rwkb model because in that sense right so so you one thing that i always found right weird for transformers but i mean it's by design is as you infer each token, you are re-computing everything up, right?

14:18That's the quadratic part. And while you're mentioning about the working memory problem, in theory with enough attention hits on it and people seem to be trying to cram more and more attention hits into the process, it could scale that way. Ignoring compute costs. Okay. Ignoring compute costs is just like a rarity, bro. That's just true as much H1N and it doesn't make sense. but RLKV was still fundamentally a neural network at its core. It ends up scaling linearly as it goes through the tokens. It will still suffer from the memory issue. So within the RLKV, we do measure two separate things. So one, we call it the perfect memory.

15:05In the model, we have only a certain amount of capacity where it can remember things perfectly, just like humans. and then beyond that, that is where it will start to discard things from its perfect memory. Right. And I felt that this was actually a lot more in line with our goals commercially and also what I felt was that it was more useful in the long run because it's cheaper compute and it could be potentially paralysable for a very long run. Right. So we're going to go into our RWQV paper in a bit. But one thing I wanted to ask, you kind of glossed over how you found it in the first place.

15:41Because you're not a researcher. You're not like, I don't imagine you're like reading papers every day or something. Until recently. Until recently. How do you find it? How did I find it? How do you know this is the one to bet on versus there's a bunch of other alternatives, right? I think what was quick, I think it was rather quick after I concluded that Transformers as it is will not scale to 10 million tokens. Okay. And so by the way, You mentioned Transformers XL. We also did an episode on Flash Attention, which helps to make part of it sublinear, at least. Yeah, but that is like way, way after I already dive into other KVs.

16:22So history-wise, at that point in time, we are talking about when 4K was the limit that everyone knew. Right. And this was last year. I mean, just to set context. Okay. Yeah. Okay. And then, yeah. So you just kind of were searching around. You found RWKV. of BKV, presumably, like, did you go straight into the Discord? Was it, like, primarily a GitHub repo? Like, what was it? Because as far as I can tell, there was no paper until maybe about two months ago. Oh, and I talked about it before the paper, right? Yes. So, you found it before they did any publicity, which is weird. It's not normal. Fair enough.

17:02So, what happened? What did you do? so what i did okay so it was basically i i believe okay so it's a mixture of things because it's like i was searching b guitar i was searching like forums other discords and also like blogs actually i was just like getting all the because everyone was just creating lists of lists right yeah yeah and i believe you also have a list of lists somewhere yeah but mine is very so i would consider myself very trad in the sense that i i would just follow the large model labs whereas the kind of list that you have to follow in order to get to something like rwkv before they've done any publicity is the non-trad like you know the kind of people that it's not working on news hermes wizard you know that like no credentials i don't even know who the hell they are but they're just working on it oh this is all for game memory and i might be hallucinating this because there was too many lists but i believe the list that actually what brought me to rwkv was that beyond OpenAI's model and beyond ChetGPT and Claudia, the two big models, right, outside of the English-speaking nations, right, a lot of the open-source models really fall flat.

18:12And that is why when you actually go through, like, lists or up for, like, doing things in other languages, RWKB actually stood out and then point. And just on the basic premise, and we're not even talking about architectural advantages, it's just the basic premise that they imported the data set in other languages in the training data yeah and was that a because I mean I imagine 99 % of your customers are English yeah was that really a driver for you it wasn't a driver you're just trying to explain it yeah that's how I landed onto like all these blocks and can you say when you say fall flat the main one that I know about is there's a tokenizer penalty for non-English yeah that's that right so like Chinese is up to Chinese or Japanese or Thai or something like it's like 16 times the number of tokens for a typical English sentence.

19:00Yeah, but even before that, right? Because, I mean, I think you understand, like, a lot of community users, they want to not use the commercial APIs. Okay. So they try to find open source models. Yes. And we'll talk about the not safe for work people. I really want, because you've actually talked to them. I have never talked to these people. But, like, when I discovered them, they are a huge community. They're extremely passionate. And they're actually good. Yeah, they're really good. They're good at this. so let's talk about them right yeah we can talk about it later yeah so so they don't want to use the commercial models and they want to use the open source model and there is a tokenizer penalty which is true but i think on the more fundamental basis right if you look through the data sets and and this is also partially informed because the way we set up our evals all evals are written in English, at least for the majority of them.

19:53And if we are racing towards building AI models, at least right now, as you see all the companies as they build their open source model, and they just want to narrowly focus on the evals, adding in a foreign data set is actually a loss. Because once you're below a certain paramount, so we're talking about the 7 and 14, the more you add that's more in line with your evals, the more it will degrade. And they just exclude it. The priority is English. Yeah, I get it. The model just fundamentally eaten. Wait, so what's the trade-off? Like, I mean, okay, so English and Chinese or, you know, there's all these other languages.

20:30What do you pick? So RWKB started with, also in context, the main person leading the RWKB project, Blink is from China. So he naturally has an interest to make sure it supports Chinese. And there are a fair amount of bilingual models, especially that are English and Chinese from the major universities in China. So we started from basically English, Chinese, Japanese, Korean. Frankly, this is large part mostly because there were fans in those communities that came on board. And then subsequently, we tried to onboard other languages as well. But these people are, again, not researchers. no money training on their home GPU lab or whatever right?

21:12Partially true but also how I see it works out for a lot of the other languages was that we have the foundation model and this is the foundation model where we just kind of said evals be them let's just make sure to include all the other languages okay and when we included the other languages right the model works for most parts for the other language. Subsequently, these individuals who wanted to use these models for their respective use cases will then fine-tune respectively. Because it's easier to fine-tune in another language for your use case than to train the language from scratch. And I think more recently, and this model is not 100 % trained yet, more recently RWKB has released what we call the world model, where we go the next step of even including all the translation data sets that we can find, even for minority languages that people end in our discord.

22:16Because the goal for them, the long-term goal for us, at least internally, is that we wanted an AI model for everyone. And everyone does not mean USA. It means the world. So there are a lot of languages in there. Is it Asia biased? Give me a sense. It's probably, no offense, it's probably still going to be US biased in terms of knowledge. Because what we're doing is still power red pyjamas for the knowledge. But in terms of language, we add all the other languages, wiki and translation set. So it's hard. I mean, we haven't fully evaluated the bias yet, but I'm quite sure that when disproportionately knowledge is still within the English universe, There's the bias there, but frankly, we are still at the stage where we can support the other languages.

23:06And I think I mentioned this. One of the interesting parallels that sometimes I have is that I can see in the Illusion forums and all that. And then we're talking about alignment. And we're talking about it in very big... Which is very keen on safety and all that, which is great. But it's not your goal as the RWKB community. yeah and and when when you talk to like members of the community came on board it's like oh i want to get this to work for korean japanese thai arabic languages and so on and so forth they just want something that worked yes they don't want it to they're not after the big model that does everything they just want something that they can play with in their language and that was very important to them yeah and these are literally just hackers doing it for personal enjoyment.

23:54Correct. Not yet for work. Yeah. Or maybe some of them for work. You don't know. We don't know. I mean, the core character AI category, there's quite a number of them using it for that. So, professionally. Professionally. Okay. As in they run character companies. Yeah. Let's call it. Should we pause here and then I'll switch to the screen? Sure, sure. Okay. Alright, so we have it pulled up. We are going to screen share for the bulk of this. So if you're listening on audio, it might be a good time to switch to the YouTube channel. So we're just going to start with an intro. What is RWKV? So RWKV is a modern recursive neural network with transformer-like level of LMM performance, which can be trained in a transformer mode.

24:38And this part has already been benchmarked against GPT-NeoX in the paper, and it has similar training performance compared to transformer models of the same data set and parent count. So specifically the GPT-NeoX model. So the key thing is that even though it's matching in performance while trading both with GPT-NeoX, it's doing all this without attention layers. And in the process, it's actually having a much substantially lower compute based on its design. And also because it's a neural network, which we will dive into later why that's substantially lower in both training and inference. and this is back to like I mentioned previously, Transformer, at least traditionally Transformer until we found out about Transformer XL and things like that, tends to scale quadratically based on the contact size and this applies not just in inference but in training and due to how this is still a neural network in its heart, even though it can train like a Transformer, it's able to do so much more efficiently and faster especially like when you hit contact size of 8K, 16K and above.

25:43And once you do like quadratic and linear, the differences start to like go crazy once you scale the numbers up. And that was the main benefits of the RRBKV model, let's say. There were a few prominent researchers when they actually reviewed through the RRBKV paper when it came out, they did highlight an important question of like, is this like evidence to literally, maybe all that really matters is that we need a large data set and a scalable model. That makes sense, obviously, to some approximation. But you are still using attention? No, we don't use attention inside. Okay, yeah. Maybe let's rewind a little bit.

26:26Oh, specifically attention as you understood it. Okay, tell us more. So we use weighted receptors. And if there's any diagrams I should pull out, let me know. Oh, okay. Okay, so we are using AFP. So this attention-free transformer, and this paper was written by Apple. What the hell is an attention-free transformer? Okay, this is unusual. Yeah, so we use the weighted retention weights, and we compute over it. And in essence, this is like the classic stacking more layers. Once you do it on top of it, you don't really need attention. Once you have enough weight, and layers stack on it. Okay. I don't know whether we want to go into the deep dive of AFT.

27:17Sure. That's interesting. I've never heard of this paper. Yeah, so this was written by Apple. And subsequently, we integrated, at least Blink, the creator of RWKB, took this and applied it to a language model and scaled it up. Right. And that is how we landed on RWKB that doesn't use attention. So sometimes within the community, we use the word light attention because what happens is that these layers and these weights will still play the role of attention. I was going to say, you end up approximating attention. Exactly. So it ends up looking at the tokens or parts of the memory and then applying it to the output.

27:58And the key benefit is that, because remember the attention model is a multi-head part. it will need to scan all the tokens back and forth. This removes that requirement. And hence, it reduced the overall compute count. I might be jumping back and forth a bit, but that's the one of the key essence of the WKV segments. And we call it light tension. And this is the part where I will disagree with the RWKV community. In some parts, I think that was a bad name. Whatever. Because it's cute. Why is it a bad name? Because when the RWKV paper came out, right, and then we talk about like we use this and we call it like attention but by design it's really nothing like your existing attention head models and it ended up like sidetracking the hacker noon debate on like one corner it's like no this is technically attention approximating attention then another group was like no this is not attention i see and and but i'm like propose better name because I have no idea what to call it.

28:59Okay. What else should people know? Maybe we can explain what RWK and V stand for. Oh, receptive with the key values. Okay. Yeah. And each of these are like actual things that you model in the code, right? Correct. So we can go into that. So which attention historically is like query key value. Okay. So do you want to jump straight into the layer architecture? Should we cover something else first? Anything good. I mean, we can, anything like high level. Okay there's a 7B, there's a 14B, there's a 4B. What are the assets or the artifacts? Okay before we go into the nitty-gritties of how the layering and everything works, on the high level currently RWKB architecturally as a model it can be, what we have already proven is that it can be scaled and trained like a transformer.

Read the full transcript

29:45How I do so we'll cover later and this can be scaled to as many parameters as we currently what we have is dominant our main models is the 7b model and the 14b model which you can find on hugging face or respectively our demos we also have on there'll be the there'll be the rwkv raven models these are also instructionally tuned okay so there's world there's raven there's music oh my god this novel what what is all this okay so we before we were The current main models is RWKB4, POW, and Raven. So POW is basically just a POW plus model. What is POW plus? I know about POW, but what is POW plus? Random data sets that the community...

30:33How many tokens were? I would just say slightly 1.1 or 1.2 times the POW. Okay. Yeah, this is not instruction tuned and stuff. Yeah, the plus one is typically all the other languages. Subsequently, Raven are the instruction tuned model. This is the current main complete models. We subsequently have... And the instruction data sets are from... Typically, GPT-4, but then we scrub it for every move or the... As a large model. Yeah, this would be the uncensored... There's some other project that's kind of doing something similar, and they call it uncensored, but really they just scrubbed it as a large-end problem.

31:14Correct, yeah. So that makes it technically breaking TOS of OpenAI, right? Yeah. Okay, but that's a later problem. Listen, frankly, let's be honest, even if we don't remove it, someone is going to remove it. I mean, so there's ways around this, which is you get clean data sets that are not GPT-4. like so the one that I typically mention is Yannick Kulture's open assistant. I believe that was included as well. Yeah obviously all these release orders are all over the place. So okay Raven World. So Raven is the instruction team model. And then subsequently the world model is a new model that we are training.

31:56It's not 100 % complete yet. We have the focus on a new tokenizer and all the languages. All the languages. All the languages that we can grab from the internet. All the wikis in all the respective languages. What do you mean when you say all languages? 100 languages. Okay, fine. So 100 languages, it wasn't really a very precise sign. We just basically, whatever the wiki tool that allows us to download the ex-wiki languages, if it works, it's in the set. If it doesn't work, skip. And all the major prominent Oscar translation sets. So as you can see, PAL, red pyjama. Alright, what is OSCQR? OSCQR is just a common term that we use in...

32:37You can just search OSCQR in Haging Face dataset, and it just means translations. Okay. So, you can find English, X pairs, all the respective pairs. Okay, yeah. So, and then all tragedy data I can find. Okay, so 70 % English, 15 % multilang, 15 % code. Is there a strong grounding for why 15 % code? No, it was just... It was already there. Yeah. The focus of the world model was not to improve everything else. It was literally that 15 % multilang. We wanted to increase... It was English and code, and then you just added multilang. Yeah, we had a bit of multilang, but we wanted to bump it up. So this is primarily English?

33:18Whatever. Okay. Yeah. What I would like is basically like a visual of like, here's all the building blocks, and here's how they combine to create all these things. So we have the RMKV architecture code. So that's the main model building block. And basically we feed it the data. Power Plus, Red Pijama, then subsequently some of the code data. For the world model, we subsequently add on top of that all the translation Oscar sets and so on. And so you're training these things. You've mentioned that you're intentionally taking a hit on evals, on traditional evals like MMU or whatever. I wouldn't say intentionally.

33:53Also to clarify, I am not training it. I'm just part of the community. The community and bling someone training. But I would say it's more like the lack of care for the evals. So the reason why we add things to the data set was never about improving evals. It's about directly in response to user feedback. It's like, oh, not good enough at this. So they're like, okay, just touch it in. Yes, literally along those lines. So take for example, right? Even for Raven and the world model, as we go through the training stages, right? We specifically ask people in other nationalities within our Discord community to test it for their language.

34:36And our rule that we set is that, our informal rule is that the only person who can decide whether this improved word model is better in Japanese or Thai or whatever it is, is a native speaker. Yeah. Where does it take place? So it's mostly in linguistics, but sometimes we do a short term in general as well. Okay, linguistics. so why don't so do you have like a appointed ambassador like you have a hundred languages yeah you just have like a czar of Japanese a czar of Thai it's not so pointed it's more of like hey this is the Japanese model please try

35:15it's not there's no the Japanese model there's one model there's a world model so if you go to world model I don't know whether it's inside here no four sorry five is you should never put five on top because 5 is fully experimental. Okay. So, under file samplersions. I see. I see. I see. So, you see there's a Japanese specific tune, Chinese tune, Arabic. Then for all the other smaller languages, we actually ask them from the base world model. Yeah. A bit itself. we feedback on that. So, we actually released previously like 10 % train, 15%, 20%. Like, as it goes through the stages and then it's like, hey, is this working?

35:54Yeah. Is it regressing? so it's like evals but real humans done by real humans and not not systematically is there a reason that you release you also so you mentioned 7b 14 14b but i see also 0.1b 0.4b 3b 1.5b like what is is that useful for people or yeah is it just for research or 0.1 and 0.4 is frankly more for research but some people do try to make use of them nothing stopping them well i mean it's extra like these are just different architectures different dimensions yeah so it's actually extra costs to you to provide these things oh but specifically for the world model what because we are trying a new tokenizer we are and and the reason why we're trying a new tokenizer is that as as i think i'm cut is that one thing that we found more like i found surprisingly frustrating existing tokenizer was that it was very English centric.

36:50And the existing tokenizer you took from GPT-Neo? Yeah. Okay. I need to backtrack a little bit, just for people who are not following along. GPT-J was the original Luther reproduction of GPT-3 and then GPT-Neo was the bigger GPT-J? Yeah. 20B? Something like that. Yeah, I do believe they have a 20B model. Okay. and there's a check for I mean for those outside of the open source space in particular for the transformer I think one thing significant for GPT-NeoX was that it was one of the major models that had everything fully documented and they like why they make this change in the architecture and so on so forth and that became like a basically reference note for all other subsequent open source models because they were the early ones that were like doing doing a good transformer model.

37:41Yeah. And at least for the large language model. So GPT-2 was actually open source. You didn't, people didn't find that useful? No, people do find, do reference as well, but it's like the code is there. And why did you do this? Oh, I see it's not documented. I see. So in that sense, was OPT from Facebook useful? Because I've heard very good things about the logbook of OPT, where they had the daily logbook and they just published that. Yeah, those were useful as well. I think one thing that NeoX had going for it, especially the elevator committee, is that it's not just logbook, it's just like, you could just go and just go, hey, why you do this?

38:24And the person who trained it will tell you. Yep, someone there, hopefully. One of them. So that's why we had the 0.1 and 0.4 models because we were just in uncharted waters here. So a lot of existing tokenizers took space as a major delimiter to detect and split. And the tokenizer we are using is actually a lot more simplified. So as these tokenizers, I mean, they scan all the text, they do a statistical model of what pairs well with what and so on and so forth. We did a similar approach, but we instead of using this token pairs well with this and should be paired with that. We just made it a trial list.

39:05So basically, we find the... Try the data structure? Yeah. So we just find the longest matching string in that matching string that we have trained inside our token list. And then we just use it as a token. It's a drastically simplified tokenizer. And it doesn't use spaces as an assumption, which I know... Which is good. Yeah. And that helps a lot of the Japanese, Chinese, and character models, they don't have spaces. And I would even argue to fair say, if you look at the really large models, like OpenAI or Cardia, tokenizers are not really a thing. I mean, in the sense that the model can work even if you tell it character by character.

39:51It's maybe inefficient. There's someone tried? I mean, there was that jailbreak where the system prompt, you put the character, then enter, enter, enter. Do you remember that geographic? No, I didn't see that one. Yeah, so you can literally, like instead of like left to right, you can literally up to down. Okay. And you're just eating tokens for every character. No, actually you're eating two because there's also the new line. And the model understood it because there's enough dumb data on the internet that it has learned how to deal with this kind of formatting. Got it. Okay. And if these models are already understanding things at the character level.

40:28Everything else is just improved compute. Because we jump the multiple tokens. Do you have any idea of your dictionary size when you use this 3D data structure? Because the typical tokenizer is like 80 ,000 tokens, dictionary size. I presume yours will be bigger. Yeah, I can't remember offhand. Our previous tokenizer is around 50 ,000. It's the new X tokenizer. Then subsequently, I believe this is around the same size. It's want to change too much on that size but we just wanted to like just change the format yeah cool i actually kind of want to establish the credentials of this thing so who is blink is rando on the internet or like because again never heard of this guy until published this is real me right and you had like i have this paper to work with but it was only published in may yeah you found this before the paper and and so i think it's very unusual for a researcher to effectively launch to the wider public without a paper and just get some kind of pretty decent community going and then publish the paper.

41:38I think a few years back, with GBT2, Transformers started to pick up Steam. And I guess the whole world is starting to think let's just abandon neural networks. So we haven't even gone into the code part, but the main reason why neural networks were bad compared to Transformer was that when you train a token, let's say you just input a token, and train a token for data samples, you have to wait for the compute to finish for that token, take the state, and then you train the next token. And we'll get into how RRWK solves that. But basically the whole work at that point just concluded, yeah, neural networks cannot scale as well.

42:14Transformer, let's just abandon it. And everyone just went in that direction. And Blink, or Bupeng is the actual name, decided basically as an individual literally at the Elutter AI forum decided that hey I think we can modify recurrent neural networks based on the Apple paper the light attention that I showed previously to make to scale this up without to make neural networks scalable and paralysable in the same way transformers work because the reason why we branch away and focus on transformers is because neural networks were slow to train it was never I mean it wasn't so much about whether was it good or bad it was just no one wants to wait a hundred years for their billion tokens to train finish even if they can throw a GPU farm at it and and that's where he started taking looking into it like how to make the neural networks trainable in parallel and and specifically RNNs yes and subsequently with the AI and I believe there was also a few others like because he was doing it very publicly there came on board to sponsor the GPU computes required.

43:25Because even though I mentioned that on large context size, it is substantially cheaper, I think, especially if you run an open source Discord forum for an AI model, it's like every day there'll be someone who thinks that they can train a 20D model on a single GPU coming in. The scale is still large. Even though it's like 1.5 or 1.10 compared to Transformer, it still needs a lot of GPU. So that's where Aether, AI, and the rest, Stability, I believe also, is involved, stepped up and donated the A100s needed to train the basic models that RWKB had. So before those models were trained, we were only having, in theory, the toy models or the smaller models that this can match Transformer.

44:14We have no idea whether it can match Transformer at that scale. Yeah, and subsequently with the larger models, the 14D models and all that, we can compare it directly with NuEx model. And that's where this paper came out. So that's the history behind it. It's like he wasn't really doing it in silence. He was doing it from Elutter. Then he branched out because this became a big project on its own. And that's where other people started coming in. Let's go. So the part where we say that RwKV is a neural network that can be scaled, can be rolled out as a transformer, right? Yep. The key thing that you want to see, right, is this diagram here.

44:56Yep. This should be in the paper. No, sorry. Yeah. Yeah. Accordingly. So what you get? So when you do inference, when you are running inference mode, ideally you should run it as a neural network. So this is a layer. so classic neural networks is that you have a state the state could be start from blank you process a token you output a state and then you rinse and repeat and then as it keeps doing the output it makes a prediction in that one thing that so subsequently for RWKB what happens here right is that we can roll out this neural network side by side and then it runs similar to transformer but the key thing here is that the states are split across the layer.

45:42So this is what we call in this diagram here specifically, this is what we call the timings and channel mix. These are operations within the layer. Depending on how you view it, you could view this as individual layers or as how we view it. We view like this collection of layers as one layer block and each layer block pass the states to its sibling subsequently down the road. As you process the next token which is a similar RNN type feature. However, the key thing is you do not need to wait for the upper layers to complete before you can go to the next token. So what happens in practice? You have to jump to the diagram like this, this graphic here.

46:25This is not 100 % how it runs behind the scene. I like it. Yeah. Whoever put time into this, kudos. I made it. So this is how you can visualize it. So the first layer is the layer norm. The layer norm doesn't... This is standard layer normalization. It doesn't need to... It just doesn't have a token and doesn't need to wait for the other layers. But if you notice, right, subsequently to the right and to the top, these tokens, these blocks, right, need to wait for the blocks on the left. And this is like, once you go past the first few tokens, right, this cascades very rapidly. Especially, like, this is only like one, two, three, four layers.

47:05Most models have like 20, 40 plus layers. and the cascading patterns are happening. And in practice, once you start cascading there, you just saturate the GPU. And that's how it starts being paralysable trained. You no longer need to train in slices like traditional RNNs. That was one of the key things. What else is the key thing? So other things is that, so I think you're familiar with LSTM, right? This is how traditional neural networks keeps things within memories. In RLKV, we have two channels. We call it the channel mix and the time mix respectively. Is there a formal definition of channel mix and time mix?

47:41Yeah. You can see the data from the respective time mix and channel mix move to the next segment. How time mix is designed, per se, was that it's how it retains, similar to LSDMs, right, where it processes a state and the input. It may decide to discard certain states and keep new things in the state. Time mix does the same thing, but with a different formula. so it replaces the LSTM in a sense and and it can decide to keep things indefinitely so this represents a long-term memory if you want to be with it that way but classically the problem with that is that it struggles with long distance correct does it have the same issue so that's that's subsequent it struggles with long distance because it also needs to keep track of both near-term memory and long-term memory so you split it up.

48:32Yeah, effectively split up. So channel mix is... Is this the perfect memory? Yeah, this is the closer to the perfect memory than it's the short term. So, so, time mix, it has trainable weights on what it decides to keep in this card. Channel mix, it has a very strong bias in it towards, like, just the next token. So, so, so, so subsequently, it was just like as, like, memories are stored in the lower layers, it just slowly shifts upwards through the channel mix. And this is the short-term memory, which at some point, as it just shifts all the way up, it will just disappear into the void. At that point, subsequently, then time mix should be retaining the longer-term memory.

49:14So we took a break for a bit, but now we're trying to cover, like, what is the big aha moment for you? And you said it was something to do with cost. Correct. So we have this chart on screen. There's literally a chart of quadratic scaling versus linear scaling in terms of GPU time spent in text generation. And you said it was at training time and at inference time. Just basically in everything that matters. All right. So, I mean, so look back to how RNN works. From a high level, we do an O1 operation on a token, create a state, O1 operation, create a state. So, this just scales linearly. You want to throw a thousand tokens at it?

49:53On inference, it just scales linearly. And subsequently, for Transformers, you take in a token, you process your first token. It may be O1 here. Subsequently, when you generate your third token, you need to compute your second and first. And then it rises. So if you do your 1 ,000 token, you need to compute back your 999 previous tokens. And as this keeps growing and growing, this is your quadratic scaling. and this is why we had this graph of the amount of cumulative GPU time that you need to spend to generate all these tokens respectively. And this is fundamentally just transformers versus neural networks.

50:34Yeah, on inference. The reason why and subsequently, neural networks did have the disadvantage of, let's say, not being able to paralyze well in training. But as I covered, RWKB kind of solved that by splitting, effectively splitting the layers, allowing you to train different parts in parallel. And like some people will go into the academic debate of like technically the second and third token is not paralysable until the first is done. But once you get into like, I can saturate a GPU land, it's just way better. It's just academic debate. We are done. And so training in essence has always, I mean, this is a bit of a transformer.

51:12On your network is I need to do an inference pass. I look at the logits. then I backprop to see what went wrong and I update the weights so the inference is the forward pass it's part of the training course as you backprop as well having needing to only look at the current cell tokens and the state instead of everything also reduce the amount of things that you need to backprop so it's just that there's so many factors involved in just reducing the overall inference and training time and that was something that appeared to me because in the long run I mean all of us want our model just run blazingly fast right?

51:48Yeah and also on minimal hardware. Oh yes Dan. Which as far as I understand you still have 14 billion parameters that's not going away. You still need the RAM to store 14 billion parameters worth of stuff. That's not going away. Yeah. Okay so RAM is unchanged. Yeah on the RAM side but the working memory is reduced. So typically you need more than 14 for transformer. I mean let's not touch quantization but in this case we don't need to keep like if you really really want to like save RAM you it is possible for you to do token by token inference so that you don't need to keep your states in history you only need to keep your current token state and your next yeah okay and yeah and that's actually like one segment of our community just purely porting RRKB to C++ based model or in the next yeah and running it on pies and stuff raspberry pies yeah it's interesting to watch there's a chart about performance and it shows that rwkb is competitive or actually better in some of the reasoning challenges which that's something i definitely would look for right like and it's it's fine if like your your speed is faster and all that but if the reasoning quality sucks then it's not a very useful language model exactly so so this is like literally us saying there's no trade-offs.

53:14Yeah, you don't lose out in that process. Okay. Big question then. Why isn't R2KV a bigger deal right now? So, one, we are not a commercial organization. Okay. This is literally the pure open source play. But you could have done the Stable Diffusion thing. Which, you know, Stable Diffusion launched. It was by a bunch of nobodies before that. It's from, like, literally split out from Luther. And, But they definitely had some hype. They definitely, like, you know, I interviewed Sharif, Shamim. The reason I ask you so many things about how did you find out how to give you, because I think the generalizable skill is how to be early in AI.

53:53Because being early in AI is very valuable. Because then you were there to see how things developed instead of, like, picking it up later like me. Anyway, so, yeah, why is it not a big deal? You don't need to be frank. Yeah. We just suck at marketing. Okay. That's fair. I mean, this is part of it. Yeah, this is part of it. So, like, maybe... But I don't think that is entirely the cause. Yeah, I'm sure you're definitely... I think the other major segment right now as well is that we were really late on the paper. Okay? Like, one of the weirdest things right now, weirdest thing right now I feel that is that RRBKB is starting to have its moment right now.

54:33Okay. Is that ever since that initial paper came out, there was ResNet, there's a... I think there's two more, there's a few more additional papers. coming up, one from Microsoft, one from other organizations that are literally exploring the whole idea, once again, of scalable neural networks. And they are citing RWKB as part of it as well. And I think for most, I think it's existingly, why switch to this model when, even though we have proven that, yes, it's scalable to 7 and 14, and that it can match transformers at similar param and training size. But all this is very academic because the community, right, the community at large, especially for the English-speaking community, right, they don't really care about this.

55:26They care about what's the best model that I can run on my computer, at least within the open source space. And even though we match in performance for things in the same data set, the keyword is same data set. Like, this benchmark is not even red pajamas yet. It's the power. And when you have models that are, be it like Falcon, being trained on much larger data set, especially for an English use case, it makes more sense to use that. I see. So there will be another paper coming that is RWKV trained on red pajama that will presumably be a larger data set. Yeah, and so on and so forth. So I think that's the...

56:06we are still in the stages of reaching that point where we train on the larger dataset. The only reason why we have a bigger outsized impact compared to the other models is frankly because half of our Discord came in not for English, it's for other languages. Yeah, that's great. And there is a definite very US and English-centric bias towards these models. And it's, to me, kind of poetic. Like there's nothing in the architecture of RDDKV that particularly bias it to be really good at other languages. It's just that as a community, you decided to prioritize it in your tokenization, in your datasets.

56:46That's it. Yeah, that's it. I would even argue that I'm more surprised that, especially on the European side of things, that we don't have more models that actually focus on even the European languages because there is like a softer jump to character Japanese and Chinese characters they're all romantic but I think back to the benchmark what excites me most still about this is that it just means that we just need to scale we just need to scale this model and we derive data to like 40B 40B 60B I mean params is one thing it's data sets and GPU time so you and I are talking offline about ideas for getting data, getting compute.

57:32So this is like a project that's ongoing. Okay, anything else for the future of Rwkb? The biggest one would be... Okay, so this is back to how, remember I said, evals doesn't hide or doesn't highlight everything. Being realistic on another weakness on Rwkb's side is that now with the rise of, let's say, 100k or 32k context size windows, transformable order, RfKB currently is trained to handle, let's say, 8 or even some people have already trained it to 16K sizes. It has, and well, as a neural network, it will happily keep going on for infinite context, man. It will just keep generating. Does it do well?

58:16The answer is no because if you didn't train it to handle that situation, and that's actually a charge. So, for example, if the prediction, the power test loss, right, it does improve over time, let's say, if we go down the context length, but this is if we train it. And what is not seen here is that if we were to, let's say, run it further, it'll just go back up. Because it was not trained to handle that. Well, it technically can run, it suffers from the longer context length, and that is the part where other TV, especially in Q &A tasks, in huge documents, like, you get closer to summarize giant documents.

58:54None of this is fundamental, it's just you need more money. yeah that's it and no there is actually a fundamental part so what what one of the things that i was doing i am actively helping within the committee right now is that we found that the existing way to scale the memory was not that efficient and we were just being realistic ourselves if we want to hit 100k we need to change this so so one thing one thing that i'm actually looking forward to right now is actually those experiments we have already we have already started scaling things to be able to handle things at transformer scale, be it the 4k, 8k, in terms of how it handles memory really well.

59:31And we found it, we want to extend it to be like 16, 32, and 64. And that is within our roadmap. And that's the exciting thing, because once we have that, it's able to handle long-term memory within those sizes, it removes what many people in the community felt, right, was the last architectural limit. Because once it's able to handle memories like context length, the same as Transformer, we know already to do all those, like, you know how existingly, like, people do, like, long conversation in Transformer, they just discard the rest and the sliding window. This is, like, the better version of the sliding window you have.

1:00:09The model can handle the sliding window perfectly. It can keep remnants behind it. Sure. And that's something that I'm really excited and invested towards because this is back to the full circle of how I came into RK. I want my model to handle 100k tokens, 4 megabytes of HTML, whatever I throw at it and be able to process it. But it'll be lossy. The later half will be lossy, but the key thing is extending the non-lossy part. And we are aiming to extend the non-lossy part. So, you have displayed today an impressive amount of knowledge just across all this stuff and you don't have a research background.

1:00:53Your advice to AI engineers getting as deep as you, who want to get as deep as you. So, I think your article articulated very well that there's going to be divisions within how we approach this. So, AI engineers, model trainers and data set curators and ML scientists. So I'll loosely define as a tree. I ignore the full stack because every company needs it. So within this tree space, there is actually a lot of ways anyone can come in without knowing anything. So let's just start with AI engineers. Don't be like, even though this whole topic, we even dived into how the layers work. We also showed how the math works.

1:01:33Frankly, for an AI engineer, you don't need it. your main thing that you needed to do was to frankly just play around with chatgbt all the alternatives be aware of the alternatives just be very mercenary swap out to cloud there if it's better for you or swap out to an open source if it's better for you and just play around the prompts learn bare prompting techniques like one shot, two shot, few shots and then from that on you can start building your agents stacking your prompts and in sequences and stuff like that and you are able to build applications that do anything in terms of the AI space and all this without knowing all this nerdy stuff with all the hard engineering because that's all you really need to actually build a product for the user, remember you are supposed to focus on making it for the user they don't care if it's RWKV or Transformer underneath the hood they just care that it helps them and I would say like Notion probably is like probably one good example of how they use it because we know underneath the hood is OpenAI but you really use it's OpenAI yeah so I obviously agree with all that let's just say that people are there already and they're just curious they want to do what you did so that's where you start going down the layers yes so the next layer the next layer you go down in is subsequently training the model from scratch fine tuning and incorporating the data set and this is this is from this is where you still do not need to know the math but you need to know like you need to have a rough sensing on how the model works and how the certain models and in this even within the open source transform certain models are better trained in certain sequences with certain learning rates and you just need to get a few of it so this is just like like the data set try it see the loss you literally did this yeah at least for RWKB and the That's a lot of work.

1:03:33Code Gen 1. Yeah, it's not a cheap work tool because you need GPUs. Okay. And that took you how long? I think Code Gen 1 alone was like six months. And then this other KB, I've been doing this for like another six months. And that is just pure experimentation. Like there's no right or wrong because like, especially if it's in a different domain. Like recently I was helping someone on the RGB Discord regarding the music generation domain. And my assumptions for learning rate and all the patterns were just completely thrown out the window because the music model just fundamentally is different in those sense.

1:04:11So that is the exciting thing is because it doesn't really have any specific rules and guidelines until you get until you travel to a certain space. It also means that you coming in is as fresh as anyone else coming in last year. it's really that kind of uncharted space for everyone and and especially as you start exploring new domains your existing your existing knowledge may actually matter because sometimes like i mean i think a few papers already covered this that's like how you train your model in certain sequences also matter like you want to train a certain set of knowledge and then and then you extend that knowledge subsequently but if you're talking about material science or genetics how am i supposed to know what is foundational knowledge what is extended knowledge i have no idea maybe you do and i'm just picking an example yeah so and the same thing for music and so on so those are things where even though you're outside the space is where you can come in just at the data set level now you want to peel off to the next layer let's just say let's just say you want to look into modifying the model the the the the the foundations of it i think one of the beauties about this current boom is that even though I did my toes early, like before the transform wave and into the early neural network phase, frankly almost everything that matters was basically in the past four years.

1:05:40Like there were a lot of things that fit in academics that were before that and you know and they were mostly dealing with models that were under a billion parameters they pretty much no longer matter and can you be more specific like okay I know I'm shooting myself in the foot because RRK is a neural network but if you're just trying to get transformers to work you don't need to know LSTM yes yeah you don't yeah there's a lot of like pre-knowledge in neural networks that is irrelevant in the transformer era and maybe some of it will have a resurgence but to get up and running is not a requirement.

1:06:21And I think this is where you could either go the very academic way of reading papers and stuff. But frankly, what I found was way more useful was Akapati. His series of videos. Serious a Hero. Yeah, that is really, really good. I think even though I read some of the papers and guides before that, it really helps that it starts from zero because you can see how it happens part by part. And even though we will not use the exact same code that he used because he re-implemented the backprop and all that, and we're just going to use Torch for that, that's where you get the aha moments on how these building blocks work and how it falls into place.

1:07:06And I had a fundamental misunderstanding of how backprop worked until I actually watched this video. Oh, really? Yeah. And I think that's the scariest and craziest thing about AI models. sometimes is that you can actually have fundamental misunderstandings but as long as you make the building blocks and you connect and okay loss is great it works yeah well so you know even the gods of the industry you know I don't know if you read the Swigulu paper so there's these like there's all alternative activation functions like there's ReLU and then there are people are always looking for different slopes and very famously the Swigulu paper had this line in there that was like yeah we don't know I don't know why this works but it works.

1:07:49Can't explain it. It literally happens here and there on the Discord too. One of the funny things that I'm doing right now in RWKV 5 experiments is that, okay, we are going to do this change, we are going to run this train, make your prediction. Will this model beat this model in this lost curve? As a game? As a betting? It's a very informal It's literally a buddy Kind of like Kind of bed But The fact that The fact that we can do This kind of bed Even though we like Understand The code It's like It just goes to show Like how often Like oh wait This didn't go What we predicted No one And And that's why Even if let's say You don't have a PhD Or So on so forth Heck Even if Math is not your special agent you're coming in as a developer.

1:08:40I'm going to say frankly, I didn't come from the research right now, the extremely math-heavy stuff is what I struggle with. What I do sometimes is I copy and paste the math into GPT-4 and ask it to explain to me. Which is good in plainer language. It's very good at that. Yeah. But the thing is, there is lots of value beyond that. One thing that I realized, and this is not specific to RWKB, this is what happens across a lot of open source models is that a lot of ML scientists when they really build this stuff the focus was more of like always get it to work it was never about getting it to work efficiently or getting the code documented or organized and stable diffusion literally went through this whole journey they had the giant they had the code and the model that worked and the community just started and engineers that came in with zero machine learning background, started picking it apart.

1:09:38It's like, no, you could replace this with this that does the exact same thing, but it's more efficient. Like, one of the major breakthroughs, for example, for GML, and this happened some time back for a bit, was that someone external from the AI community went in and implemented memory mapping. Yes, I saw that. I forget her name, but yeah. Justine.law is her URL. Yeah. And she didn't come in as an AI expert. She came in as a software engineer. Yeah. And these are all just very, very straightforward. You know, in her world, this is normal. Whereas for the researchers, they will be like, I don't want to.

1:10:19Wait, what is memory map? Yeah, exactly. Yeah, and there are a lot of things. Like, one of the jokes that I have right now is that every month there is a research ML scientist that's rediscovering the number 32. Why? Because be it like someone in the community writing the inference code, because GPUs, especially QDA GPUs, tends to work really well when they align to the batch size of multiples of 32. And if you've been in the gaming industry, especially when you write shader code, this is like well-known, just given knowledge. And people are just constantly rediscovering, oh, maybe if I just adjust my data set, my data size to fit this batch size suddenly I get 10 % improvement and yeah and I was like these are things that once again because they were so focused on just making it work that they won't know outside space and that's why I would say right if anything right now is the best time is that you don't know AI to have people from different background come in because your contribution could be from dataset level, how to train the knowledge, to shader code, to hack, how to memory map, how to cache data.

1:11:37There's so many gaps. Cool. Great. So yeah, thanks so much for being very willing to get on and talk with no prep. We did some prep, but it's very unusual podcast episode, but I really enjoyed it. We literally just met yesterday in Singapore. But I know you've been on the Discord for a while and I can tell you like you're very serious about all this. I think it's very unusual for someone like you have a job, but this is like a second job, essentially. But you are really enthusiastic and passionate about it. And I think that's very rare. And I don't want to encourage more people to do it. And so thanks for sharing.

1:12:09Thanks for having me here.

From the publisher

The AI Engineer Summit Expo has been announced, presented by AutoGPT (and future guest Toran Bruce-Richards!) Stay tuned for more updates on the Summit livestream and Latent Space University.

This post was on HN for 10 hours.

What comes after the Transformer? This is one of the Top 10 Open Challenges in LLM Research that has been the talk of the AI community this month. Jon Frankle (friend of the show!) has an ongoing bet with Sasha Rush on whether Attention is All You Need, and the most significant challenger to emerge this year has been RWKV - Receptance Weighted Key Value models, which revive the RNN for GPT-class LLMs, inspired by a 2021 paper on Attention Free Transformers from Apple (surprise!).

What this means practically is that RWKV models tend to scale in all directions (both in training and inference) much better than Transformers-based open source models:

While remaining competitive on standard reasoning benchmarks:

swyx was recently in Singapore for meetings with AI government and industry folks, and grabbed 2 hours with RWKV committee member Eugene Cheah for a deep dive, the full recording of which is now up on Latent Space TV:

Today we release both the 2hr video and an edited 1hr audio version, to cater to the different audiences and provide “ablation opportunities” on RWKV interest level.

The Eleuther Mafia?

The RWKV project is notable not merely because of the credible challenge to the Transformers dominance. It is also a distributed, international, mostly uncredentialed community reminiscent of early 2020s Eleuther AI:

* Primarily Discord, pseudonymous, GPU-poor volunteer community somehow coordinating enough to train >10B, OPT/BLOOM-competitive models

* Being driven by the needs of its community, it is extremely polyglot (e.g. English, Chinese, Japanese, Arabic) not because it needs to beat some benchmarks, but because its users want it to be for their own needs.

* “Open Source” in both the good and the bad way - properly Apache 2.0 licensed (not “open but restricted”), yet trained on data taken from commercially compromised sources like the Pile (where Shawn Presser’s Books3 dataset has been recently taken down) and Alpaca (taking from Steven Tey’s ShareGPT which is technically against OpenAI TOS)

The threadboi class has loved tracking the diffusion of Transformers paper authors out into the industry:

But perhaps the underdog version of this is tracking the emerging Eleuther AI mafia:

It will be fascinating to see how both Eleuther and Eleuther alums fare as they build out the future of both LLMs and open source AI.

Audio Version Timestamps

assisted by smol-podcaster. Different timestamps vs the 2hr YouTube

* [00:05:35] Eugene's path into AI at UIlicious

* [00:07:33] Tokenizer penalty and data efficiency of Transformers

* [00:08:02] Using Salesforce CodeGen

* [00:10:17] The limitations of Transformers for handling large context sizes

* [00:13:17] RWKV compute costs compared to Transformers

* [00:16:06] How Eugene found RWKV early

* [00:18:52] RWKV's focus on supporting many languages, not just English

* [00:21:24] Using the RWKV model for fine-tuning for specific languages

* [00:24:45] What is RWKV?

* [00:33:46] Overview of the different RWKV models like World, Raven, Novel

* [00:41:34] Background of Blink, the creator of RWKV

* [00:49:55] The linear vs quadratic scaling of RWKV vs Transformers

* [00:53:29] RWKV matching Transformer performance on reasoning tasks

* [00:54:31] The community's lack of marketing for RWKV

* [00:57:00] The English-language bias in AI models

* [01:00:33] Plans to improve RWKV's memory and context handling

* [01:03:10] Advice for AI engineers wanting to get more technical knowledge

Show Notes

Companies/Organizations:

* RWKV - HF blog, paper, docs, GitHub, Huggingface

* Raven 14B (finetuned on Alpaca+ShareGPT+...) Demo

* World 7B (supports 100+ world languages) Demo

* How RWKV works in 100 LOC, RWKV overview

* EleutherAI - Decentralized open source AI research group

* Stability AI - Creators of Stable Diffusion

* Conjecture - Spun off from EleutherAI

People:

* Eugene Chia - CTO of UIlicious, member of RWKV committee (GitHub, Twitter)

* Blink/Bo Peng - Creator of RWKV architecture

* Quentin Anthony - our Latent Space pod on Eleuther, coauthor on RWKV

* Sharif Shameem - our Latent Space pod on being early to Stable Diffusion

* Tri Dao - our Latent Space pod on FlashAttention making Attention subquadratic

* Linus Lee - our Latent Space pod in NYC

* Jonathan Frankle - our Latent Space pod about Transformers longevity

* Chris Re - Genius at Stanford working on state-space models

* Andrej Karpathy - Zero to Hero series

* Justine Tunney ("Justine.lol") - mmap trick

Models/Papers:

* Top 10 Open Challenges in LLM Research

* Retentive Network: A Successor to Transformer for Large Language Models

* GPT-NeoX - Open source replica of GPT-3 by EleutherAI

* Salesforce CodeGen and CodeGen 2

* Attention Free Transformers paper

* The Pile

* RedPajama dataset

* Monarch Mixer - Revisiting BERT, Without Attention or MLPs

Misc Notes

RWKV is not without known weaknesses - Transformers do well in reasoning because they are expressive in the forward pass, yet the RWKV docs already note that it is sensitive to prompt formatting and poor at lookback tasks. We also asked pointed questions about RWKV’s challenges in the full podcast.



Get full access to Latent.Space at www.latent.space/subscribe

More from Latent Space: The AI Engineer Podcast

All 247 episodes
RWKV: Reinventing RNNs for the Transformer Era — with Eugene Cheah of UIliciousLatent Space: The AI Engineer Podcast · 1 h 12 min
Listen in VO