Data synthesis for SOTA LLMs

6 Feb 2024 · 47 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Practical AI Podcast Episode Summary

Episode Title

Data Synthesis for SOTA LLMs

Episode Description In this episode, Karan Malhotra from Nous Research discusses the origins of Nous, a collective of LLM researchers, and the techniques behind their state-of-the-art data synthesis methods that contribute to the popular Hermes family of models. The conversation delves into fine-tuning strategies and the effectiveness of data synthesis in enhancing model performance.

---

Key Participants

  • Karan Malhotra - Co-founder and researcher at Nous Research
  • Chris Benson - Tech strategist at Lockheed Martin
  • Daniel Whitenack - CEO and founder of Prediction Guard

Links

  • [Nous on Hugging Face](https://huggingface.co/NousResearch)
  • [Nous Research Website](https://nousresearch.com/)

---

Key Concepts Discussed

  1. Origins of Nous Research
  2. Nous started as a distributed collective of AI enthusiasts and researchers, initially motivated by the desire to work with open-source models.
  3. The organization evolved from a community of open-source practitioners engaged in fine-tuning and model training.
  1. The Hermes Models
  2. The Hermes series is a product of collaborative efforts, inspired by earlier models like GPT-2, Llama, and GPT-J6B.
  3. The models were developed using synthetic data derived from more powerful models (e.g., GPT-4) to improve performance.
  1. Synthetic Data
  2. Synthetic data is generated by language models and is used to fine-tune other models, thus allowing smaller models to compete with larger models.
  3. The process of distillation is crucial, which involves compressing knowledge from larger models into smaller, more accessible forms.
  1. Model Training and Fine-tuning
  2. The importance of hyperparameters in training models is emphasized.
  3. Karan suggests ignoring conventional limits on training duration; if the model isn't overfitting, continue training to maximize performance.
  1. Community and Collaboration
  2. The growth of Nous Research has been organic, driven by community involvement and volunteer contributions.
  3. The team uses platforms like Discord to coordinate efforts and foster collaboration across different projects and specializations.
  1. Future Directions
  2. Nous Research aims to focus on locality and offline capabilities, allowing users to run models independently without relying on cloud services.
  3. Despite transitioning into a corporation, their ethos remains rooted in open-source principles, seeking to enhance the community rather than restrict it.

---

Key Takeaways

  • Data synthesis is a powerful method that enhances the performance of smaller models by using distilled data from larger models.
  • Collaboration within the AI community has led to significant advancements in model development and accessibility.
  • Future developments will prioritize user autonomy in AI usage, enabling more individuals to harness the power of language models on their terms.

---

Conclusion The episode encapsulates the spirit of innovation and community in the AI field, showcasing how collaborative efforts can lead to the creation of impactful tools and technologies. Karan Malhotra’s insights into Nous Research not only highlight the organization's achievements but also set a vision for a more open and user-centric future in AI.

For further information and to join the ongoing discussions, visit [Practical AI](https://changelog.zulipchat.com/#narrow/stream/456003-practicalai).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:28Welcome to Practical AI. link in the show notes. Thank you to our partners at Fly.io. Launch your app close to your users. Find out how at Fly.io.

0:43Welcome to another episode of Practical AI. This is Daniel Whitenack. I am the CEO and founder at Prediction Guard, and I'm joined as always by my co-host, Chris Benson, who is a tech strategist at Lockheed Martin. How are you doing, Chris? Doing great today. It was nice seeing you a few days ago in person. In the flesh. In the flesh. Yeah, that was great. I think you posted a picture on LinkedIn. So if anybody doesn't know what we look like and has some crazy reason to want to know, there's a smiling mug of us on Daniel's profile. Yes, yes. And the reason we met is I was on a client visit on site and we were prototyping out some stuff like chat over your docs and natural language to SQL stuff and all sorts of things with PredictionGuard.

1:33And one of the models that we were using was from Noose Research. And that works out great because we have Curran Mahotra here, who is from Noose Research, co-founder and researcher there. So welcome. Glad to have you, Curran. Hey, all. Thanks for having me. I'm extremely excited to chat with you guys. Yeah, like I said, I'm a huge, well, this is our first time meeting, but I feel like we're already friends because I've had so much of my own benefit and interaction in working with models from Noose Research. a lot of amazing models that you've posted on Hugging Face and research that you're doing.

2:13I'm wondering if you could just give us a little bit of a background about Noose specifically and kind of how you came together as researchers and started, to me, from the sidelines, it seemed like, oh, all of a sudden there's these amazing models on Hugging Face and I don't know who these people are, these Noose research people, but they're amazing. So give us a little bit of the backstory there. Absolutely. Yeah. So just as a general overview, we are one part like open source research organization. We put these models out for free. We put a lot of research out for free, some data sets so people can build on top of these open models.

2:51On the other hand, we're very recently a company as well, a C-Corp. So we've been working pretty hard after getting some seed funding on building together some exciting stuff. I won't go too into during the overview point, but we're continuing to do our open source research and development and release of models indefinitely. The way we started is very interesting, and it would be pretty out of nowhere to the outside for sure. It was extremely fast for us. We're a collective of people who have been playing around in the open source language model space for a while, ranging from like GPT-2 release to Llama release to like the first Transformers paper.

3:33We've got people from various eras of Gen AI of when they came in. And for myself, it was GPT-2. I stumbled upon a CoLab notebook and started fine tuning, made some Edgar Allan Poe and Lovecraft tunes. I've done the same. That's awesome. And we just got pulled into this world of look at these next token predictors that are just managing to smatter together the most wonderful and amazing stories that slowly turn into a deeper and deeper dive of, well, how can I use this for learning information? How can I learn to use this for production and automation, right? It's evolved over time. For us, we started off just working with different open source collectives, actually.

4:16Once OpenAI kind of released GPT-3 and had closed sourced it. We were used to open source GPT-2. We were like, oh man, what are we going to do? How are we going to continue to play with the level of customization and interactivity that we had with GPT-2? Then Eleuther had released GPT-J6B, the cobalt AI community, this community of people who tune models and inference models started to pop up, I think around 2020, 2021, in the face of this. So a lot of us started to have places to centralize and play with these models. We got to contribute and learn how to become better open source AI developers, et cetera.

4:56Eventually, there was a need for more concrete organizations to do this kind of focused work on the creation of these models. We were stuck with like, okay, architectures for a while, like Pythia, but thanks to Meta, you know, we wouldn't be here without Meta. I'll say that first and foremost, like - The great Llama. Yeah. Yeah. Like prior to Llama, right? Like everyone's like, oh, Facebook evil, like my data, et cetera. And here we are, like, they are kind of like the shepherds of this new era of the open source AI movement. So when Llama came out, there was a paper that came out called Alpaca by Stanford lab, right?

5:36And this was about distilling data from bigger models like GPT-3, ChatGPT, GPT-4, and being able to train smaller models on that distilled synthetic data, something they call the instruction data. So that alpaca format really opened up the playing field for everybody to start making these instruct style models, these actual four prod use style models. So there was an idea I had in my head of, well, the Alpaca guys are using only GPT 3.5 outputs. What if I only generated GPT 4 outputs? It'll be a little expensive, but you'll probably get a better model out of it than Alpaca. At the same time that I was looking at this, there was a guy on Twitter named Technium who had just started putting together his own synthetic data set based off Alpaca and the GPT 4 only as well.

6:26So I was working with a group at the time called Open Assistant under Lion. They're a really big nonprofit. And while I was working on that, we had some GPUs. They were cool with us using towards the development of new models. So I reached out to Technium and I said, hey, I have a little bit of compute. You have GPT-4 data in the same format. I have GPT-4 data in the same format. Let's train a model. So we trained a model called GPT-4 x Vicuna. this model was on the vicuna fine-tune we fine-tuned to fine-tune basically the vicuna model was a alpaca style fine-tune and we tried our data set on top of it it was good it was okay then we thought you know we'll probably get a better result if we just train on the base llama model and the resulting model was the very first hermes model gotcha the og the og and and that's kind of how it started to come together was we both had a data thesis on use GPT-4 only and follow Alpaca.

7:28And we trained on Lama and we got Hermes. And we didn't know what benchmarks were. We didn't know anything about any of this stuff. We just made a model and it got a ton of attention. We put it out under this name, Noose Research. Noose comes from the Greek word for intellect. We thought it would be a good name for an AI company. But it was just a place for, you know, fun projects and fine tunes and stuff. It was just a name we were using for our collaboration. And people started swarming and asking, you know, what's news research? Like, what's this sudden, like, mystical, like, open source organization that, like, put out this, like, best model?

8:07And we're like, best model? Like, we just, you know, we just tried something. It was really organic. And it got to the point that people started telling us, you must have trained on the benchmarks. These are doing too well. And we were like, what's benchmarks? We're not really coming from an academic place as much as from an enthusiast that became so committed that it became our life. It became our day-to-day. So from there, people started to ask us, can I join news research? Now, there wasn't a news research to join. It was just two guys, right? What ended up happening was we formed a private Discord server, and we thought there's a lot of people who range from somebody who's like 16, 17 years old, savant on Twitter, hasn't even been to college yet, insane at Transformer stuff, to mid-30s, working a really, really good fang-esque job, and just wants to really create and let loose.

9:05that was another class of volunteer. And then you have, you know, older gentleman who has already exited a company or something who has just been playing with code for a while and wants to jump in and hang out. So we ended up being this really eclectic group. You know, we don't know what your name is. We don't know what your race is. We don't know your gender or anything. It's just discord profile picture, Twitter profile picture. Right. So we came together, grew to about like 40 people all working together on various different projects like Hermes Tunes, Data Synthesis, the Capybara series, Context Length Extension, etc.

9:39And just from this kind of interaction between Twitter and Discord and bringing people in that we thought were cool, we ended up becoming what people would call an open source research org. Yeah, you sort of stumbled into creating this amazing research organization, which is ruling the world, which is awesome. It's what OpenAI might have been. Oh, well, yeah. That's really sweet. Thank you, guys. Yeah, and I love it. It's so cool to hear that story and that background. And I see in my own sort of little snapshots here and there, I'm connecting that in my mind over the past couple of years as I've seen you all post different models and that sort of thing.

10:24This is something we've definitely touched on on the show before, but some of our listeners might not kind of fully grasp when you say this sort of like synthetic data sets that you were focused on in this alpaca format. Could you kind of explain a little bit, like we've talked a lot about fine tuning and, you know, preference tuning and RLHF and different things, but what does it specifically mean that like you would take synthetic data? What does that mean in your case? And like, why does that result in something good in fine tuning an open model? People might think, oh, this is synthetic data.

11:02Why should I expect it to like be any good? So could you kind of help explain that subject a little bit? Yeah, absolutely. So, I mean, out of context, synthetic is like as meaningless as like artificial, right? Data is data. But in this case, it's referring to a particular class of data that's been generated by another language model or another AI, another diffusion model, et cetera, that can actually be used to further train models. Now you might say, why would you want to do something like that? How is it helpful? What was important to us is we were all GPU poor, right? We were all running on laptops or maybe a 3090, maybe a 4090.

11:39As individuals, we don't have data centers. So training or even tuning a large model in the early days, like 70 billion parameters, something like that, was just unfeasible for us. And knowing that GPT-3 is like something like 175 billion parameters and 3.5 and 4 can only go up from there. The question became, how can we make these small 7 billion parameter models even compete with these massive, massive ones? These ones that I want to run offline, these ones that I might want to run on an edge device, on a phone, on a drone, etc. Right? Like, how can I make them even useful? So there's two things to talk about here.

12:18One is synthetic data and the other is distillation, right? So synthetic data is just referring to like any kind of data that's created by a model in this case. And the reason that's useful is in particular distillation. So if I told you to go study comp sci for 10 years, for example, and put in that massive time investment and really focus on general programming, and then I told you, you know, now it's time for you to learn about AI and transformers and stuff and put you through all the math prerequisites, et cetera. Like you're going to come out with like a really strong foundation of how to do the work.

12:56But the problem is you've put in a massive time investment. Now, if I take that guy who spent 10 years doing engineering, then another five years doing AI, and I ask him, Hey, can you teach somebody like just really important, like compressed tidbits that'll help them just get up and running to do the work? That's data distillation, right? That's knowledge distillation. So you look at these big models, like a CLOD or a 70B model or GPT-4, and you can see they're amazing. They're brilliant at everything. They have a bunch of high-quality data they're trained on, and they have a bunch of low-quality data they're trained on that they can interact with and express in a high-quality form.

13:36So instead of me having to read a massive 10 pager for why some chemical reaction or some like tax based process, whatever you want it to be like, instead of reading a massive document on that, and then feeding that to a language model, we can just have that really smart model that already understands it really well, compress that information into an instruction, or into a conversation until like two sentences, three sentences, five sentences, like half a page. And we can just train a much smaller model on that compressed information, and it will learn the compressed information, you know, to the degree that a language model learns something, you know, not perfectly.

14:18But because of that, what the Alpaca guys did was they generated a bunch of seed tasks from GPT 3.5 on various different domains and topics and created these kind of compressed instructions with instruction, an input question from the user, and then an answer. So the instruction could be like, given the following math equation, explain step by step why this is the answer. And then the input is the equation, which is your question. And then the output is the compressed answer. So all of that we can take as one sample in the data set, and we can make hundreds of thousands or millions of samples like that of various different domains and various different tasks.

14:59So the alpaca guys did this, less than 100K examples, I believe. And they trained the llama models on these, and they found massive boosts to performance that this distilled information, like a human, successfully compresses and transfers over. So when I saw that, and then independently when Technium saw that, and then independently when many others saw that, we were like, this is so intuitive. This is exactly how I've learned anything by just going on Discord and Twitter and bothering people to give me the compressed bit of how I do something. We should try doing this with even higher quality models than 3.5.

15:34So we created, I can't remember the exact number at the moment, but at least 50 ,000, maybe 100 ,000 examples originally for Hermes one like this, just using GPT-4. And then we trained on that and ended up getting performance that was extremely, extremely like massive boost compared to the other models that were not trained using this kind of method. So without these giants that have already established themselves in the space, we wouldn't be here. Like without OpenAI, without meta, like we literally wouldn't have the model and the data to do the kind of work that we did to make Herbys. What it allowed for us is like for local models to finally be like comprehensible and for us to finally have like offline capabilities to kind of take the good stuff from something like GPT-4 or something else and make it uncensored.

16:28So it still has all this understanding of all these topics, but it doesn't have all that RLHF inside it necessarily that safetyizes it so that when people utilize the model has all this intelligence, but it has more freedom of thought to kind of converse with you on topics that open AI may reject. Gotcha. One of the things I was curious about as you were going through that was a few episodes back, Daniel and I were kind of talking about the effect of model licensing, you know, on the community and the different kind of licensing concerns that were coming out from whether it be, you know, meta open AI, you name the organization.

17:03Is that ever a challenge for you since you're kind of using those to get started in terms of the inputs? Has that been a concern or do you anticipate it being a concern? I think that, of course, generally, US international regulation on this stuff is evolving. The conversation is evolving very much. So naturally, there's like, you have to keep it top of mind. You have to think about these kind of things. But thankfully, because all of our model releases are like open source and we don't profit from them. If somebody goes off and creates a product using our model, good for them, but we don't necessarily take on the liability or that worry of saying, hey, we're going to sell you this model that was created with GPT-4 outputs.

17:46We actually actively try to stay away from doing that. But because the data distillation paradigm is so effective, if a model comes out that's better than GPT-4 and it's open source and I can use it locally. and in their TOS it says you can use this to make a commercial model, then we can apply the same techniques that we've been preparing and researching and understanding from these closed models and use it there. So right now we don't stand to or try to or have any plans to profit from using any of these outputs. We're not about that because we want to be careful and respectful of these model creators and these companies.

18:25But that being said, we're learning all these techniques and developing all these techniques that will be useful for when that time comes and for when that's available, especially with the advent of something like Mistral. If we do distillation from a Mistral model, like Mistral Medium or something like that, that's completely, from my understanding, barring their TOS saying otherwise, but I believe it doesn't. It's completely okay in that situation for us to create models like this that can be used commercially, et cetera. Regarding the TOS stuff though, as much as we err on the side of caution, I'd find it hard to see a company enforce their TOS when these larger models are likely trained on not all copyright-free stuff.

19:14I'd find it hard-pressed to believe that these closed source companies, their models are totally copyright-free and totally copyright-clean. So if some other company that was feeling a little more rambunctious than ourselves was to say, you know, we are going to commercially release on this, I imagine it'd be difficult for them to be come after without the other group opening their books. And there's actually a pretty interesting interaction that happened regarding this between Google and OpenAI, if you guys are familiar. here. So yeah, I saw this interesting picture the other day. It was like the interesting web of AI and it was like how Microsoft, Google, OpenAI, it's like on one side there's the ones and it shows how they're connected to the other ones.

20:00It's like this visualization and like how many of them overlap in these strange ways between like whether it's Together or Mistral or Meta, Google, Microsoft, OpenAI is sort of very interesting web of connections that probably make some of these things rather difficult. Leave it for the lawyers to sort out. Yeah. Yeah. That's the thing is like, we can look at an example, right? Like you hear that phrase, like good artists copy, great artists steal, right? Like, so the data distillers, we're copying, right? Like we're just distilling this information. Like we're trying to like make our models more like those.

20:38And we don't really plan to commercialize. We're just doing it for free for everyone. But the great artists are, you know, Google, you know, like you look at Bard and it tells you, you know, I was made by OpenAI. Now it's fine for our open source model to say I was made by OpenAI because we're very transparent about this is trained on GPT outputs. But when Bard violates the TOS with a paid product. Bold. Yeah, that sounds like I was trained by OpenAI, right? You think that OpenAI would come after this multi-billion dollar company like immediately, right? Instead, you see a tweet from, first you see Google deny it.

21:12Then you see a tweet from Sam Altman, which was something along the lines of, I'm paraphrasing here, something along the lines of, I'm not mad that they trained on our outputs. I'm mad that they lied about it. And I'm sitting there like, okay, you're mad about this, but aren't you going to pursue the legal action in your terms of services? No, no. Because everyone would have to open their books up too. That being said, I don't condone the commercial use of that kind of stuff. Like making a paid model from GPT-4 outputs, I wouldn't advise anyone sell a model made with them just because we want to respect people's TOS and stuff.

21:51They worked hard and spent billions to make this stuff or hundreds of millions, however much they spent. But there is certainly room for hypocrisy in that realm of the large corps. So that's my thoughts on the licensing stuff. And that's definitely my own individual thoughts. We're a pretty decentralized collective at Noose, so you'll find people with all sorts of opinions all over the place. And as a company, we don't hold any view whatsoever on that. Yeah. I'm wondering, maybe this gets a little bit to the distributed nature of this, but I know that there's sort of various collections of what the Noose Research Group has done over time.

22:33You mentioned Hermes, but then there's these other kind of categories of things too, like the yarn models, capybara, puffin, obsidian, just looking over the hugging face now. I'm wondering if you could just give us, from your perspective, a little bit of a map of these different things and how people might categorize the different collections of what Noose has done. I definitely want to talk about the future things and ongoing things as well. But as it stands now, what are the major categories of what the collective has invested their time in over time? Certainly, certainly. So within the stuff that's viewable on Hugging Face, at least, we've got the Hermes series, of which, like I told you guys, the initial story of how it went down.

23:20But from there, Technium kept going. I haven't personally had any interaction with the Hermes model since the initial. From there, Tech just continued to create more and more synthetic data, collect from more and more sources, use more and more open data sets. And he's just got the, I guess, award-winning data thesis. The guy really knows how to go about curating and synthesizing good data. So Technium, it's his baby, the Hermes project. So everything you've seen since is really his work and anyone who has kind of collaborated with him. But almost like you can't call it anything a solo project because of the open data sets we use to like everything is built on the shoulders of giants and the shoulders of each other as little people.

24:03But tech really has helmed the Hermes initiative so far. I think that's our most popular model series. And he released the open Hermes as well, because we had some data in the original Hermes that we never released publicly. And we wanted to make that kind of an option for everybody. So that's Hermes. Still follows the same kind of philosophy of synthetic data. And it now uses the chat ML format instead of the alpaca format is what we kind of upgraded to. Then you've got a Capybara and Puffin, which are both done by a volunteer and, you know, OG member LDJ, who you may be familiar with, Luigi Daniel Jr.

24:40So the Capybara series was using an amplify instruct method, this novel method that LDJ had worked on alongside another one of our researchers, J. So LDJ and J can get confusing, but the two of them worked on the Capybara series, created the dataset, trained the models. And then Puffin was the idea of using hand-picked smaller samples from some of our larger data sets to make sleek data sets for an easy tune and see how that works kind of in the spirit of the Lima paper, where they just used a few examples to get really good results. Those are really the popular tunes using synthetic data for general use.

25:23Yarn is this novel context length extension method at the time of creation by Emozilla, also known as Jeffrey Cannell, and Bowen Peng, also known as Block 97, alongside Enrico Chipotle and Eleuther AI. So what happened there was these guys were already looking into contests like the extension for a while. And when we kind of came under the noose banner to do the work, it opened up a little bit of resources from compute sponsorships. It opened up a more centralized place for them to be able to do that collaboration. I had no hand in the yarn models whatsoever. And that's the exciting thing is everyone really gets to work in their own spheres and their own kind of autonomous circles.

26:10And then we just check in and see, you know, how's the research going? How's it coming along? Because we really work with people that we heavily believe in and we believe in their idea. So if we don't already have an idea, we're kind of just say, please freely create because we brought you in because what we will freely create will push forth our agenda anyway. So I think those are our big model releases and series that we have available. Outside of that, we have a bunch of stuff on our GitHub as well. Stuff that's being worked on, stuff that hasn't necessarily come out yet. There's a lot of that.

26:44So I got a question for you as a follow-up. It's pretty fascinating, the story that you've been telling us here because of that kind of organic, you know, creation of the organization or collective. And I'm wondering as you've done that and you kind of went through and talked about the different model groups and kind of talked about, you know, the owners or spiritual owners, if you will, of each of those families, how do the different members of the collective interact to kind of share? Like how do you each push each other along or share information or give ideas so that cross-family efforts can kind of benefit from the overall collective.

27:21And as you said, now a C-Corp and you guys are more organized at this point. So what kind of culture has developed around those communications and learnings? Yeah, absolutely. I mean, when it started, it was just like a small Discord, maybe like 10 people. From there, like we kind of created more channels as people wanted to work on more things. and we had initially split up into like three, four different topics or sectors that people could assign themselves to. One being data synthesis, of course, so we can kind of find new novel methods and formats for distillation and the creation of synthetic data.

Read the full transcript

27:54One being training, like people who are just like really good at training hyperparam stuff and people who will come up with new architectures and new techniques. Another being agents, a group of people who want to actually try to build tools and do autonomous work with this stuff. And then we had this one category that it was a prediction for the future of simulation. So we had people that were very interested in kind of bringing this stuff into simulation, into Unity, into kind of seeing how all these things came together. And it was interesting because the training built on the data synthesis, the agents build on the training, and then the sim would build on the agents.

28:29It was kind of the idea. So everybody needed to work together because all those things are so intrinsically connected, but people would have specializations on kind of where in that workflow they wanted to work. We didn't end up doing a lot on the sim side of things. Now, recently, there's a lot more interest because we have a lot more capability generally as the AI community does. But as we've grown to, we went to 40 people, it was fine. Now we've gone to like 5 ,000 people in the Discord. It's a little unwieldy there. So what we do is we kind of tier people in, you come into the Discord, you can see maybe two channels, and then we'll give people a developer role.

29:08We don't really let people select their own roles because we want to make sure we can sort through people we know to let them through. And even as we do open source research, a lot of it is unreleased, and we want to make sure that it's protected before release. So we create this developer role so people can then see way more channels of just general development and development conversation. And from there, as we see contributors who have started to do more work or show more passion towards contributing to news in a particular field or who have some reputation or some portfolio in a particular field, then we'll assign them one of those roles.

29:47And that will open up the family of channels relating to those roles and our current projects surrounding that role. So data synthesis projects, agent projects, training projects, etc so we kind of just tier it out so people can interact and people have been around for a while or people we consider fellows or part of the cohort they can usually see pretty much everything so they're pretty effective in serving as coordinators for the cross communication between these different channels and groups and even if something has like a particular someone has a particular role or some channel has a particular role it's supposed to be a part of Like it's still discord and we're still very chill.

30:25So like people will still work on like various different overlaps inside of just one channel as well.

30:43If you're listening, you know that artificial intelligence is revolutionizing the way we produce information, changing society, culture, politics, the economy, but it's also created a world of AI-generated content, including deepfakes. So how can we tell what's real online? Read, write, own, building the next era of the internet, a new book from entrepreneur and investor Chris Dixon explores one possible solution to the internet's authenticity problem, blockchains. From AI that tracks its source material to generative programs that compensate rather than cannibalize creators. Read, write, own is a call to action for a more open, transparent, and democratic internet, one that opens the black box of AI, tracks the origins we see online, and much more.

31:32This is our chance to reimagine world-changing technologies to build the internet we want, not the one we inherited. Order your copy of Read, Write, Own today or go to readwriteown.com to learn more.

31:57I have a selfish question, which now that this is one of the advantages of doing the podcast, I get to talk to all the amazing people doing amazing things and learn from them. But I'm wondering as a person who is also trying to fine tune some models, either just for my own enjoyment and learning, but also fine-tuning models for specific tasks and in specific customer use cases and that sort of thing. There's a lot of people out there, I think many of our listeners who are thinking like, since you being part of this collective have worked for, you know, since the sort of dawn of these many, you know, the proliferation of fine-tunes from Llama and etc.

32:41And as you've seen all that, as you're doing more and more fine-tunes now, as you're looking towards the future. Do you have any kind of good advice or things to keep in mind for all those like fine tuners out there that are thinking about grabbing something off of Hugging Face, creating their own versions of these models? Maybe they have their own ideas about a specific take on a model. Any general tips that you found to be really useful over time or like pitfalls that you'd like to highlight? Yeah, I mean, I can try to think of a few off the top of head. I'll say that hyperparameters are really important and it's important to try to get that right.

33:22It's going to vary from model to model, but a lot of the time, some people think hyperparams don't really matter as much to obsess over. And some people think it's like a secret sauce as well. So I'd say like try to do a lot of research into like good hyperparams, a good learning rate. Like I'd also say like, I could be totally wrong about this as I am not the trainer of Hermes today or a lot of these models, but something I personally believe in a lot is ignore people telling you to only train for X amount of time. If you're not overfitting, just keep going if you can. If you have the compute, keep training and keep going.

33:59Train for more tokens, more epochs. That's something I heavily believe in. In terms of trainers to use, there's a lot of people who make their own scripts for specialty stuff. And there's, of course, you can just use Hugging Face, but the library we use is called Axolotl, A-X-O-L-O-T-L, like the animal, by Cassius Wing Leon of the Open Access Collective. We think Axolotl is probably the best general purpose trainer for LORAs, QLORAs, fine tunes, et cetera. It, like any open source repository, is buggy and stuff you're going to have to work out but it's in my opinion probably the easiest and most effective trainer to use for like pretty much any model architecture available right now so i definitely point everybody towards axolotl awesome yeah that's super useful we'll share some links in uh in our show notes as well so people make sure and check that stuff out another kind of interesting question As you see, I think we saw these waves of models that came out maybe around synthetic data, fine tunes, or other types of fine tunes.

35:11I see this interesting sort of thing happening over the past however many months, not that long in the scheme of things, but in the AI world, maybe a while. where we're kind of now, like there's a lot of interesting approaches more so than just fine tunes, but like mixture of experts and merging and of course multimodal stuff coming out. Now I see news kind of dabbling in that. You don't have to answer for the whole collective, but as there's so many of these things coming out and different approaches, what are some of the things within that, doesn't have to be one of those, but what are some of the things on your mind kind of moving forward or on Noose's mind kind of more generally?

35:54Sure. I'll try to go from like simple to complex on the kind of stuff. That sounds great. I think that definitely just like straight up instruction tuning is great. There's other ways to tune, like the Evol instruct method. I would advise people to try to create new instruction methodologies that allow us to make even better formatted data. People don't spend enough time trying to create new instruct formats. and we've definitely been swamped with not doing that as well. So I think towards the general community, it's a really easy place to get started. You don't need to really know how to code so much as think about how a human might more effectively phrase something or format something and kind of remix from there.

36:35I think that's like probably the easiest place to start. Then there's a model merging, right? Model merging is great. You can just like take two models and Frankenstein them together to question mark results. You know, you got to just try and see what happens and feel it out. Then from there, I would say there's stuff like DPO. There's RLHF, DPO, like this kind of rewards things that can let you like enable rejections or create censorship or put some kind of general concept or attitude towards the model. We found that to be pretty effective with the latest Noose Hermes Mixtral DPO. It seems like people really like it and prefer it over just the SFT.

37:16So that's another thing that I'd heavily recommend. From there, we get a little more complex. We have some reward model stuff we're working on that I won't speak to just yet outside of saying we're working on it that we think is going to be pretty big for reasoning boosts. Of course, there's techniques like chain of thought and tree of thought for multi-step prompting, creating data sets even out of that for any of these purposes I've already mentioned is going to be really effective. Now, to stuff that maybe not everybody can, actually, a lot of people would already be able to do this. There's something that we like to call over at Noose Activations Hacking, where you're kind of messing with the way that a model, I'm trying to think about how to say this in the most layman's terms, you're trying to mess with how a model like generally vibes about something.

38:05So rather than just doing a system prompt or something like that, you can actually like change the model vectors to kind of be like more political about something, less political about something, more terse, more specific, and it has far more effect and control over a model than a system prompt. It's basically like a system prompt that like tells it to embody certain characteristics, but it's not something you can really jailbreak or get around as far as my testing has shown, certainly not as easily as a system prompt. We have no problem jailbreaking even the most censored closed models today.

38:39It can be done by anybody with the right words, right? But this activation stuff, it really creates a bit more of a robustness and fidelity to the concepts that you're trying to tell it to embody. There's a few more I'm trying to think of that would be useful for people. One thing is soft prompting. It's not really around anymore. It used to be pretty big during the GPT-J, like pre-Llama days, when the Cobalt AI guys really pioneered the use of it in the open source community. But a soft prompt basically takes like massive prompt and compresses it down to like way less tokens. So you can give your model like a huge prompt, a huge system prompt or huge amount of information and use like way less tokens.

39:22So soft prompting is cool. It's not going to be too difficult to like update it for like llama mistral like today's architectures it's just like nobody has really done it that i've seen so you know to the community if you guys do that please share um that's actually much easier than the activation stuff i think and then finally probably the hardest unsolved is like uh sampling methods like today we use like top k top p like you know nucleus sampling, et cetera, whatever. There's better ways to pick tokens, for sure. There's better ways to judge the value of tokens, for sure. Everyone has been too concerned with higher levels to get that low and do whatever the magic math is that I can't do that would enable some steering and some, even beyond steering, alternative sampling paradigms.

40:17And I think that would probably bring the biggest change and transformation to literally all models, regardless of the tune, regardless of the architecture, et cetera, get pulled off. So really looking forward to something like that happening in the space. That was a lot of really good advice that you have there. I was sitting there trying to take notes while you're talking through it and everything going, wait, but he said that too. And he said that too. No, the really good answer there. Thank you for that. As we're starting to wind up here, I wanted to ask you, I know about as we're recording this is looks like it was just over three weeks ago, about four weeks ago when we release this episode, you guys announced your$5.2 million seed financing round.

41:01So congratulations on that. That was pretty amazing. Thank you. And I'm kind of wondering, so like you've kind of started with this kind of fairytale story of kind of organically building from the ground up, you know, yourself, you connect with somebody else, a few other people join, you get to thousands of people contributing you fine and really producing amazing work. And then you're incorporating and now you got the seed round coming. Where does that lead you? It's kind of a sky's the limit kind of scenario, it seems, you know, that now that you're, you're kind of launching and, you know, on that, you know, as a corporation, as you said, where can you go from here?

41:41What do you anticipate over the next couple of years or even several years out? You know, what's the vision. What do you want to achieve? You've come a long way so far. What's next? AGI. No, I'm just kidding. I believe you if you said it, actually. I mean, like, you know, someone will do it. And then you'll distill the knowledge. Then we'll distill and then you'll run the AGI on your Neuralink, on your contact lens or something. But for us, like, there's a huge focus on locality. There's a huge focus on offline. There's a huge focus on take the power back, run the model yourself, do everything at home.

42:21That's big for us. And at the same time, of course, we believe in scale. But there's this idea that there's so much unsolved at the small model size. Why don't we do that before we go to a trillion params? Because we can scale those realizations. But for us, there's certainly a transformation and change in attitude and in pressures from going from pure open source volunteer to as well having kind of this more corporate branch get created as well. But that being said, it's been pretty consistent, our ethos and our motivation for why we do this. And like you said, it really was organic in the sense that we're a product of the times.

42:57We're a product of the atmosphere of the AI community. People have said nice things like you guys are setting the trend. And it's not really true so much as the truth is we are one of many embodiments of the sentiment that the community has and that the world has, we think. There's more than one noose research in this world. You know, there's alignment labs, there's Pygmalion, there's COBOL, there's people who have been around before us, people who will come along the way, people who have already formed since we have. And there's lots of people who have kind of embodied the noose research ethos.

43:29And it's not really just our ethos as much as the overall community's ethos. There are people who have come before us, people who will come along the way, who do very, very similar style of work as us, this kind of open work. And I think that's got everything to do with the fact that this is what the people want. We're just the everyman, just like everybody else. We're not billionaires or super all-X Facebook or anything like that. We're just a bunch of people who really, really care about this, who want to see everyone have access to language models, everyone be able to automate their lives, everyone be able to push their understanding of any topic to the next level.

44:12And our work as we become an organization that's looking to be a company and create revenue, et cetera, we won't let it tamper or hinder any of the open source work we do. In fact, we wanted to empower all of that work because we believe that the tools and the developments and services that we will be providing as a corporation will only serve to better feed the entire open source community. We're not really looking to suddenly make like a closed Hermes or something like that. We're more looking to create tools and do research that makes your open Hermes far more effective, far better. And, you know, good enough that you may want to pay for that tool.

44:58It sounds like something I would pay for. That's for sure. Yeah, it's super inspiring. I really appreciate you taking time, Curran, to talk with us. I thoroughly enjoyed this because I am such a fan of everything you all are doing and the community that you've built. So thank you for saying true to that culture and what you're doing. And I'm really looking forward to seeing what happens in the future and where things head. And I hope that we can talk again and have Noose back on the show in a year when, of course, everything will be different in the AI world. And I'm sure you'll still be doing interesting things.

45:35So yeah, you're always welcome back on the show. Thank you so much. It's been a pleasure to chat with you guys. Thanks for being so candid. I'm glad we were able to kind of push our message forth more. And thanks for the validation you and the community have given us to keep doing this great work. All right. Thanks. We'll talk soon. See ya.

45:59that is practical ai for this week thanks for listening subscribe now if you haven't yet head to practicalai.fm for all the ways and don't forget to check out our fresh changelog beats the dance party album is on spotify apple music and the rest there's a link in the show notes for you. Thanks once again to our partners at fly.io to our beat freaking residents, Breakmaster Cylinder, and to you for listening. That's all for now. We'll talk to you again next time.

From the publisher

Nous Research has been pumping out some of the best open access LLMs using SOTA data synthesis techniques. Their Hermes family of models is incredibly popular! In this episode, Karan from Nous talks about the origins of Nous as a distributed collective of LLM researchers. We also get into fine-tuning strategies and why data synthesis works so well.

Join the discussion

Changelog++ members save 2 minutes on this episode because they made the ads disappear. Join today!

Sponsors:

  • Read Write Own – Read, Write, Own: Building the Next Era of the Internet—a new book from entrepreneur and investor Chris Dixon—explores one possible solution to the internet’s authenticity problem: Blockchains. From AI that tracks its source material to generative programs that compensate—rather than cannibalize—creators. It’s a call to action for a more open, transparent, and democratic internet. One that opens the black box of AI, tracks the origins we see online, and much more. Order your copy of Read, Write, Own today at readwriteown.com
  • Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs. 

Featuring:

Show Notes:

Something missing or broken? PRs welcome!

More from Practical AI

All 157 episodes
Data synthesis for SOTA LLMsPractical AI · 47 min
Listen in VO