In short
Eye On A.I. Podcast Episode Notes
Episode Title
#147 Yilun Du: AI Debates, Reinforcement Learning, & The Power of Generative Models
Episode Overview
- Host: Craig S. Smith
- Guest: Yilun Du, PhD student at MIT EECS
- Topic: Discussion on generative AI, reinforcement learning, AI feedback techniques, and the future of AI in robotics.
Key Themes and Discussions
- Introduction to Yilun Du
- Background in AI research at OpenAI, FAIR, and Google Deepmind.
- Focus on generative models, decision-making, robot learning, and embodied agents.
- Reinforcement Learning with Human Feedback (RLHF)
- Definition: A method used to improve AI models by incorporating human feedback into the reinforcement learning process.
- Challenges:
- Inefficiency due to reliance on human feedback.
- Inconsistencies in human ratings can lead to unreliable model training.
- Reinforcement Learning with AI Feedback (RLAIF)
- Concept: An extension of RLHF that utilizes AI agents to provide feedback.
- Advantages:
- Reduces the need for human involvement, thus eliminating bottlenecks.
- Can potentially provide a more scalable solution for improving AI models.
- Multi-Agent Debate Method
- Mechanism: Multiple AI agents generate responses and debate their accuracy, leading to improved outputs.
- Benefits:
- Encourages reasoning and critical evaluation among AI agents.
- Helps identify inconsistencies and enhances overall model accuracy.
- Applications and Future of AI Training
- Discussion on the gap between open-source and proprietary models.
- The importance of increased computational resources for advancing AI capabilities.
- Vision of decentralized AI systems that can operate independently and collaboratively.
- Challenges in Robotics
- Integration of AI with robotics for creating intelligent physical agents.
- Recognition that while AI algorithms are improving, the physical implementation in robotics still faces significant hurdles.
- Need for better hardware and more intelligent algorithms to drive advancements.
Key Takeaways
- RLHF vs. RLAIF: While RLHF relies on human feedback, RLAIF seeks to leverage AI's ability to critique itself, enhancing efficiency and scalability.
- Multi-Agent Approach: Engaging multiple AI systems in a debate format can yield better answers by allowing models to challenge and refine each other's responses.
- Future Directions: The discussion highlights a vision for a more decentralized and collaborative AI ecosystem, particularly in the robotics field, where physical interaction with the world is critical.
Conclusion The episode emphasizes the ongoing evolution in AI technologies, particularly concerning training methodologies and the integration of AI into robotics. Yilun Du's research offers insights into innovative methods like RLAIF and multi-agent debates, indicating a promising future for both AI and robotics as these fields continue to converge.
Sponsors
- Crusoe Cloud: Provides high-performance cloud computing solutions optimized for AI applications.
- Celonis: Specializes in process mining to enhance business workflows with AI.
For more information, visit
- [Crusoe Cloud](https://crusoecloud.com/)
- [Celonis](https://www.celonis.com/?utm_source=eyeonai&utm_medium=referral&utm_campaign=ai/)
---
Listen to the Full Episode You can find the full podcast episode and transcript at [Eye on A.I.](https://eye-on-ai.com).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00One of the fundamental limitations of ROHF is that you don't actually know if the model has learned the thing you want to teach it right? You have no guarantees that the model has learned what you have taught it. Some of the most successful companies right now are all in charge of basically collecting data. And I guess they're not covered as much in the news, but I feel like that seems like a huge amount of people. If you have this AI feedback and it can give you answers for everything, right? You would hope that the coverage is actually a lot more than the coverage you could get from human feedback.
0:32Hi, I'm Craig Smith, and this is Eye on AI. Large language models are notoriously inaccurate, stringing together coherent sentences that often have little grounding in reality. To fix this, companies like OpenAI deploy armies of humans to try and nudge LLMs toward more accurate responses using reinforcement learning, a process called reinforcement learning with human feedback, or RLHF. This is obviously an inefficient and labor-intensive way to tune the world's most powerful AI models, and researchers are looking for ways to automate the task. In this episode, MIT PhD student Ilun Du talks about new techniques for improving large language models through reinforcement learning with AI feedback, or RLAIF.
1:32Ilun explains how RLAIF circumvents the need for expensive and inconsistent human ratings by having AI agents debate and critique each other's responses. He shares insights from his research using multi-agent debate to enhance reasoning and accuracy, and we discuss the challenges of accessing proprietary models. I hope you find the conversation as interesting as I did. Hi, this episode is sponsored by Salonis, the global leader in process mining. AI has landed and enterprises are adapting, giving customers slick experiences and the technology to deliver. The road feels long, but you're closer than you think.
2:22You see, your business processes run through many systems, creating data at every step. Solonis reconstructs this data to generate process intelligence, a common business language. With process intelligence, AI knows how your business flows across every department, every system, and every process. With AI solutions powered by Solonis, enterprises get faster, more accurate insights, a new level of automation, and a step change in productivity, performance, and customer satisfaction. Process intelligence is the missing piece in the AI-enabled tech stack. Search Celonis, C-E-L-O-N-I-S, to find out more.
3:12This episode is sponsored by Crusoe Cloud. High-performance cloud computing or low environmental impact, Crusoe Cloud was built because the innovations of the future need both. Crusoe Cloud is a scalable, clean, high-performance cloud optimized for AI and HPC workloads and powered by wasted, stranded, or clean energy. Crusoe offers virtualized compute and storage solutions for a range of applications, including generative AI, computational biology, and rendering. Visit CrusoeCloud.com, that's C-R-U-S-O-E-C-L-O-U-D dot C-O-M, to see what climate-aligned computing can do for your business. I'm currently a PhD student at the Computer Science and AI Laboratory.
4:10So I've been working on generative models for the last five years, especially on their application to construct intelligent physical agents or robots. So, I guess another name for generative models recently has been generative AI. So, I've been working in this space for the last five years. And yeah, mostly I've been working a lot on diffusion models for quite a few years. And then recently also a reasonable amount on language models. Yeah. So, I'm interested in RLAIF, reinforcement learning with AI feedback. And I guess my first question, the reason I'm interested is because that's how OpenAI and the other large language model providers are trying to nudge their models into good behavior to stop hallucinating and things like that, which to me seems like a very unsophisticated method to use armies of workers around the world sort of, you know, scoring responses.
5:29So RLAIF is a natural extension or solution. but am I wrong that reinforcement learning with human feedback really began as a strategy in robotics? So not quite. Actually, in 2018, I was working on OpenAI. So that's when they just started actually this ROHF. So there was a paper between OpenAI and DeepMind. I think this was in 2018. That was, which was, I guess RL itself is a field that has been going on for like many, many decades. Deep reinforcement learning is also a thing that actually has been going on for even like for like two decades. Actually, even before DeepMind's DQN, there were quite a few works on deep reinforcement learning.
6:26But anyways, this ROHF idea at the time in 2018 was that RO agents would often exploit your environment, and they would exploit parts of the environment that you really would not want it to. So essentially, maybe you wanted to complete this race course, but it finds that it gets a lot of points by looping around the ship. So instead of completing the race course, it will end up looping around the ship an infinite number of times. So initially, the goal was that by having humans give feedback to our agents when playing these interactive games, you can prevent the agents from doing prevent the agents from like doing very odd behaviors.
7:09And then this evolved, I think around 2018 was when the first GPT model came out. So then there is the safety team at OpenAI. And I remember there were several people who were interested in the idea that you could use this idea of RL supervision to train the language models to summarize or to have humor. Because it seemed like these are things that are very hard to teach the language model if you just gave it a bunch of internet text. So I guess that might have been the start of this idea. And then afterwards, there were a variety of papers. and then eventually, I guess, like, I guess ChatGPT was a version of this, like, RL from human feedback.
7:54Right. And just for those listeners who don't understand how it works, can you describe the RLHF, how it works in the sort of process? and when you were working on it in OpenAI, was that before the GPT models or was that on early iterations of the GPT models or were you using, what were you using RLHF for? So I personally was not using RLHF. At the time when I was at OpenAI, it was a very small company, I think around 40 people. And then maybe, so we knew, I knew about all the research that was going on at the time. So this was right when the first GPT-1 model came out or before that. So this was very early on.
8:48The focus of the company at the time was RL. And then there was focus on this game called Dota that was going on. And the idea of RLHF is that your model will generate different outputs, and then you'll have people rate them. People will say, this is very desirable. This is not desirable. uh and then you'll basically and then this gives the model supervision on uh on like if you should generate this output um more often or less often and and as i said uh since these language models have blown up and hallucination uh has is such an obvious problem the there people have been using rlhf and and it's actually uh to to guide the the models into better behavior uh two things about that fascinate me one is that uh you're only guiding you're only kind of nudging them the model uh and and uh is there some metric for how many times you need to push the model to stop it from doing something you're pushing it away from but also uh as this has as llms have exploded uh there's like this new industry of human reviewers uh around the world and do you have any sense of how many people are involved in rlhf for large language models at this point i mean the human reviewers?
10:24Yeah. So I guess going on the first question. So yeah, I think that's one of the fundamental limitations of ROHF is that you don't actually know if the model has learned the thing you want to teach it, right? You have no guarantees that the model has learned what you have taught it. And then there's a variety of papers that have shown that you can take any of these large language models and just by adding some additional prompts, it's very easy to make them elicit behavior that you would have hoped that you had made the models forget. So yeah, a big issue with ROHF is that the reward function just nudges the model towards doing something you want it to do.
11:04It doesn't necessarily guarantee that your model has learned this. On the other hand, though, you could say that human learning, also, as a teacher, you kind nudge a student with, this is the response you want, this is not the response you want, but you don't necessarily know if the person has actually learned the thing you wanted to learn. Maybe they just have copied the equations you've given it. But yeah, so along the first question, I think that is an issue with ROHF. Along the second question, how many people are rating these models? That is actually something I don't really know very well.
11:43My understanding is it's a lot. So I I think based off what I've heard, some of the most successful companies right now are all in charge of basically collecting data. And I guess they're not covered as much in the news, but I feel like that seems like a huge amount of people. And I think the job is very, I guess you could say boring and maybe scarring also because you see all this stuff on the internet and then your goal is just to label things. Yeah. And the, yeah, I guess that's right. I mean, I hadn't thought of it in that regard. It is just labeling, right? It's like supervised learning labeling.
12:29It's just for a different purpose. And when you said some of the most successful companies, you're talking about data prop companies or? Yeah, like data, like, yeah, like, yeah, uh data companies that are focused on gathering data yeah yeah um so uh this bringing ai into the picture seems like a natural solution uh and can you talk about how you do that how do you get an ai to rate uh a response uh from from a model yeah so So instead of, so like when you do ROHF, what happens is your model will generate different outputs and then people will just say, this is good. This is not good. So in a very similar vein, you can imagine that you have two language models.
13:29One language model generates the answer and then you give the other language model that may be the ground truth answer and the question. And now it can like basically assess how good the generating answer is. like first, does it match the ground truth answer? And then they can like check the reasoning, like does the reasoning make sense to the language model? So that's a very simplistic version, you can imagine of ROI for an AI feedback. And then there's like a variety of different ways to make it much more sophisticated. So one way that I've been working on a lot is this idea of using multi-agent debate.
14:05So what happens is now, let's say I give you a question. what happens is you have multiple instances of your language model generate different possible answers to the question. So it's kind of like when I give you, when you're trying to solve a math problem or you're trying to answer a question, you have different thoughts in your mind. And each of these voices in your mind kind of talk to each other. So each of these language models talk to each other, try to, is this answer consistent with my first thought? Is it consistent with my second thought? And then you can imagine that you have these language models debate with each other.
14:37And then the answer you get at the end of this debate should be a more accurate answer than the original answer you have. So you can imagine that now this gets you better data, right? And then you can again use this data to retrain your model. The core idea is basically you can use the model to get better data. You generate from your model and then you can impose some structure on this generation process to encourage the model to do a bit more reasoning or causal chain of events or stuff like this. And then And at the end, the answer you get is an improved version of what the model would originally generate.
15:12So it gives you some improvement operation. Yeah. And is there... So in the first, the more simplistic scenario where you're giving a model the ground truth in the prompt, is that... you do that in the prompt or you're you're you're connecting a vector database with with uh you know verifiable or ground truth data how do you yeah i think probably the simplest way is to just give it as a prompt uh so so you can ask the first language model here's a question generated answer and then the second language model uh you give the original question again and then you give the generated answer and then you say here's the ground truth answer and then and then you can instruct the language model uh reflect and like uh describe what is correct versus not correct about the answers so walk me through that let's say uh you ask a model what's two plus two and the model answer is five and then you have another model yeah so so yeah so let's say the first model generates two plus two is equal to five so what you do now is you give the second model you say here's uh so you give it the same question what is two plus two and you and then you give it round truth answer for and then you give it answer generated by agent uh two plus two equals five uh and then you can say uh uh can you check if this answer is correct and provide feedback on what portions were correct or what portions were incorrect.
16:56So one thing you could do is you could feed the intermediate generations of the language model because the language model generates the answer step by step. And you can give it the first sentence and say, is this consistent with the right answer? You can give it the second sentence and say, is that consistent? And you can get much richer annotations than just, this is correct, this is not correct, which is in principle something you could also do with people, but it's very tedious for people to do this. So in practice, most people, most of the time, the ROHF feedback that you get is just like holistic over the entire output.
17:30Yeah. So in these LLMs talking to each other, you don't have a human operator that's asking one LLM and then takes that answer and asks another LLM this is all happening within the software between the two LLMs. Is that right? Or is there some human oversight required? Yeah. So in this debate speech feature that I was talking about earlier, yeah, there is no human oversight. Well, I think one big thing about RO for an AI feedback is this idea that you do not want humans to be their bottleneck. Because LLMs or all of these models, they can be accelerated very quickly on hardware accelerators. But people operate at a specific speed.
18:28They're expensive. Maybe they're inaccurate. So yeah, the goal would be that completely automate this as much as possible. Right. But where does the LLM, the reviewing LLM, get the ground truth? data from a vector database or uh yeah well i think there are two two two ideas right so one idea is that you yeah you could give that uh you could give the uh like the ground troops llm the correct answer right like if you if there are a set of problems that you want to teach the llm to be more accurate on right you could uh you could just uh generate a bunch of questions get a bunch of answers right also and then like use the cell of them to like critique your answers.
19:13I think the thing that people a lot of people want is I assume that you do not have the ground truth answers. So even if you do not have the ground truth answers right like let's say let's say I want to like write a proof for like a mathematical theorem. I don't know the ground truth answer but I can like check to see if someone else's answer is incorrect. So I can can like clearfully look at the first step and say, well, this step is logical or this step is not logical. I can check the second step. So one thought about this, like ARL for AI feedback is in the same way to like how people, you can like prove mathematical proofs by sequentially like just checking and making sure everything is correct to your knowledge.
19:55You can imagine each of these LLMs like check each other and see is every single step the way correct? And if every single step is rational and you get the answer you want at the end. Right. And again, even if you don't know what the ground choice answer is, we still know this is a satisfactory answer, or this is not a satisfactory answer. So you can imagine it's a way to search through the space of possible solutions. So having one LLM generate many different things and having the other one prune the ones that are probably wrong. And if you prune all the bad ones, then you get the right answer.
20:29And then if you have a model that can always generate the right answer, right, it's a more intelligent model than the one that's not, that's just generating all possible answers. And this, I mean, I understand how that could work for a logical proof or mathematical equation. But how does that apply to a more general and less precise questions and answers? I mean, if I ask an LLM, what's the history of this town? And it comes back with a hallucination that's not correct. How does the other LLM know to check that first LLM's answer? Yeah, so I think this is where, so I guess I talked a bit about how, like, in some of my work on ROEIF, RL for an AI feedback, I use this idea of having multiple models.
21:35So I think if you only have a single answer, right? So like, if I ask you like, when was this person born? And then the LLM just makes up some number. There's no way to verify if it's correct or not. But the thing about LLMs is they can generate multiple answers for a question, right? So if I ask the LLM, when was this person A born? And then it gives me 1876, and then it tells me it's 1896, and then it tells me it's 1895. Now you have multiple different answers, right? And now your critic can say, well, clearly something's amiss here, because none of these answers are consistent with each other.
22:13So I think when you introduce this idea of not just having one model generate, but have multiple instances of the model generate different solutions, then you can kind of verify issues with the causality or the chain of reasoning of language models. You can identify things that are logically inconsistent by just having multiple generations come out. yeah uh so by having multiple generations of or multiple agents debating as you said at the beginning is that is that what you mean yeah that's what i was thinking and and i mean uh this is all happening in some vector data vector space right it's it's not happening in actual
23:08language on on a screen and that's then being fed to the other agent how does how mechanically how does it work yeah so actually uh so far for the work that i've been working on debate uh you actually have the language models generate language outputs uh explicitly and then each of the language models takes in these language outputs from other models and then uh it gives feedback on that so So it's purely in the language space. And probably it could be more efficient to be able to do it in some type of vector space. But actually there's some desirable properties having it in language. If you can have all these agents reasoning language, then we can kind of understand what the agents are doing also.
23:52It's a hard constraint that forces the model to be kind of reasonable in a sense, because if you do it in some abstract space, maybe you as a person, we have no idea what the model is doing. And it might also be finding all types of like cheating ways to like communicate. So like having it, having all of this self-improvement be directly in language, I think is a desirable thing also. And so in your research, are you getting printouts of these conversations between AI agents? How do you monitor what the conversation or the debate is between agents. Yeah, you can actually, yeah. So you can actually see the entire printout of the conversation.
24:38So like it starts with like, you give it a question and then model and then agent A will say, based off the order of operations, I think it's X, Y, Z. And then here's the final answer. And then agent T will say, based off this other thing, I think the answer is A, W, A, B, C. And then each of these agents will respond to each other. So like one agent will say, well, actually, I think the second agent is wrong. I think it should be this. And the second agent will be like, oh, no, I am correct. And then you can have this multiple rounds of conversation. And then essentially at the end of this conversation, each agent will be like, based off what I've seen of the other agents, I think the answer is ABC.
25:16And then you just take this final answer. And it will converge eventually to one answer? Or do you end up with LLMs that will never agree? Yeah, so most of the time, you end up converging to the same answer, or like converging to one answer. I think this might be due to the fact that the models... So the models are trained with... These models have been trained to converse with people, right? And they've been trained to converse with people in such a way that they are not very stubborn. If a person's constantly suggests an answer, they will respond to it. So what we notice is when you try to repurpose this mechanism to have multiple models talk to each other, right, then the models are very definite to other models.
26:04So it's almost all the time they will converge to one answer. And then you check that answer against ground truth or against sort of... Yeah, basically we just check against ground truth and we'll see how good it is. Yeah. And so what are the, presumably you have metrics of how different strategies have improved response accuracy or something like that. But how do you measure that? So, yeah, so there are a set of existing benchmarks. And then what you can do is, and roughly they're in the style of, here are a set of math problems, and I know the ground truth answer. So basically, you take the final generation from all the agents, and then just compare the accuracy in which that final answer matches the real answer is one thing you can do.
27:00And in a similar sense, for factually correct things, you can ask the model to generate all these facts, right? And then you can get ground truth facts about the subject of interest. And then you can basically check the accuracy in which the language model generated things matches the ground truth. So all of these problems, to evaluate this approach, you can always come up with a database of answers and check those. I mean, the hope would be that you should you would always have these benchmarks that you know the answers for. So you can like concretely monitor monitor monitor the progress. But then ideally, you'd want to do it on all types of subjects where you don't have very clear ways to monitor how good it is also.
27:43Yeah, because and again, is this in the in the fine tuning stage that you're doing this? where you're trying to improve the accuracy of one model by having it debate with these other models? Or is this an architecture for, as a consumer, you know, I interface with a language model on my screen, but in fact in the background there's this debate going on between multiple models that then converge on an answer before supplying me with the answer i mean is this a training strategy or is this actually a new way of of building large language models yeah i mean so so in the in the paper we wrote about this we primarily focused on this application where you finish training this model.
28:51And then this is a way to get better answers from the language model. But I mean, I think one thing that I've been very excited about is this idea that this is also an improvement operator, right? During training time also, this is a way to get like RL for AI feedback. And like I've been recently talking to several companies that I've also been like, I guess I've been training the large language models and they're also like thinking, like exploring one of these strategies. uh so yeah so i mean i i mean it's it's i think it could be either right in general i think anything that you can do for like any technique for ai for an uh like rl from ai feedback each of these each of these strategies gives you some mechanism for how good an answer is right so you can always use that as a test time thing also as well as a training thing i like there i think they're like two flips of the same coin like i don't think one is very different than the other yeah uh but but But if you're using this strategy to train the model, that's one cost.
29:53But if you're building it into the model so that the model is going through this process to produce every inference, it seems to me that would be extremely costly. actually on one on one side if it's it's if it's built into the inference every time you ask a question before you get the output there is this debate going on in the background that would increase the cost right uh yeah but but then on the on the on the other hand if you're using it for training and this is something about rlhf also or or you know trying to nudge these models you you do it for example on uh you know questions where there's a very clear ground ground truth answer but that doesn't mean the model will necessarily generalize to all questions right uh at training time uh well yeah yeah so so i think yeah so definitely it's a case if you use that at test time the computational cost is much more uh i feel like it's always it always seems like a thing though that if you have a thing that improves performance the actual computational cost is always like not the thing to care about the most because that cost always decreases year on year.
31:23So I feel like it's always been the case that every year on year, things get much, much, much cheaper and cheaper to do. So if you have something that can fundamentally improve the performance of your system, it seems like worrying about cost is not the most important thing because people always find ways to make things faster. But yeah, on the latter point, So, yeah, so if you, well, ideally, any type of RL from AI feedback, you would be able to apply on arbitrary prompts, hopefully, because ideally you construct this interface that all the models can just autonomously talk to each other and generate answers.
32:04And as long as the models can autonomously generate answers, you can give it all types of prompts. One issue with having people rate the accuracy of something is it's really hard to rate the accuracy of not so well-known people. Because you would have to search for ages on Google and maybe to verify one passage that's generated, maybe you need to spend two hours. So it's not very scalable. And then for lots of prompts, maybe there's just no easy way to do it. But then if you have this AI feedback and it can give you answers for everything, right you would hope that the coverage is actually a lot more than the coverage you could get from human feedback right in the training phase you mean uh yes in the training phase in the training phase because that that's an issue uh right when when i mean the hope in in rlhf is that the model will eventually, it'll learn improved behavior.
33:06It'll learn to reason more carefully or to learn to double check itself or something in the background so that it then can generalize that behavior to any question that it's asked. But that's not necessarily the case. Is that right? Yeah, I mean, yeah, I guess with all these things, you can never know, like you can never guarantee, like this model will like see every possible thing, right? These models have like billions or trillions of parameters. So it's hard to guarantee, but the hope would be that you get more and more coverage, like the more scalable your automated critique is, the more possible problems you could cover.
33:51And hopefully, if you cover all possible problems, maybe it will be very difficult to elicit odd behavior for the models yeah um and on the on the uh improvement on the measured uh improvement um the um is there have you i mean how do you measure that so against the benchmarks like uh before rlai f uh you know you get a 60 score score by, and that's another question, how do you score, I guess, on a benchmark? Yeah, so you have a 60 % accuracy. And then after RLEIF, you've moved it up to 80%. I mean, what are the numbers that you're working with in terms of the improvements that you're seeing with these strategies?
Read the full transcript
34:47Yeah, so I think right now, yeah. So what you would see is like maybe on this benchmark, it used to be 60%. Now it goes to 80%. So yeah, you would want to see these consistent improvements across a wide variety of benchmarks. I think there is a question where let's say the models get really good, which I think they're very, very far away from at the moment. But if they were to get very, very good, and maybe they do perfectly at all these human benchmarks, then how do you assess the performance of these models. And again, I think there's a nice idea from this multi-agent debate kind of thing, because you can use something like an MMR system.
35:26So for all these chess and goal-playing bots, they are much better than people, but we can still quantify how much better they are by pitting these agents against each other. So if the agent is able to beat another agent 100 % of the time, then it's probably a much more intelligent agent than the original one. So the hope would be that if you use this type of MOTA agent competition, as your improvement operator, you can also use the ease at which the language model is able to generate solutions to convince another agent as a measure of how intelligent it is. So if the model gets more and more intelligent, ideally it would be much better, it would maybe blow away the other model.
36:10You can speak so eloquently or so correctly that you just cannot, you will always compete the other one. And then maybe that could be a metric for improvement at that point. But I think that's very speculative. I think the models right now are very, very far away from being human level. And probably my guess is in the next year or so, or at least upcoming several years, probably we will have plenty of benchmarks to see numbers go up a little. yeah but but in your research how how much have the has the uh performance improved with these strategies i mean do you see a two percent improvement or yeah i see like a 20 improvement which is a decent amount uh and i think uh so we wrote this paper in may uh and i think there's been several follow-ups uh on this paper and i think some of the large companies i think for example, Anthropic might also be exploring debate.
37:06But yeah, it seems like there's a boost. It doesn't seem like, well, we don't have access to the large language model, so we actually cannot do this full-on training at a large scale. So we don't see these new capabilities or anything like that. That would be the hope if it could. But yeah, right now we see this boost in these raw numbers. Yeah. And if you don't have access to these proprietary models, what language models are you using in your research? Yeah. So we have API access to these models as everyone else does. So this way we would propose to do this RLAI feedback, it doesn't require the weights of the network.
37:54we can directly... It's more of a test time thing, like we were talking about earlier, where you can still use this to improve your performance at test time. But yeah, the big issue is we can't fine tune the model or use that to improve the performance. Yeah, I think that's a bit of an unfortunate thing at the time, I think, because there is a pretty large gap at the moment with all these open source models and the ones that are being used commercially. So it feels like likely the pace of research or the pace of improvement could be improved a lot if the weights were made available to everyone.
38:34But the issue is that I think there's so much money and incentive on commercial interest on this side that that's unlikely to happen. But yeah, because... But what about using LAMA 2, for example, where you have the weights? Yeah, so I think the issue with LAMA 2 is still it's pretty poor, actually, at the moment. So lots of people talk about how, oh, it's really good and stuff like this. But it still is like, so some papers claim that if you use LAMA 2, you do lots of fine-tuning, it can match GPT 3.5 performance. What we find is that, yeah, it can match GPT 3.5 performance on one task, which is a task they spent the last month optimizing for, and it's far worse on every other task.
39:25And then even GPT 3.5, for this debate stuff and a lot of these AI feedback stuff, GPT 4 is much better at this than 3.5 is. So it feels like if you're stuck with things that are much worse than 3.5, it's very hard to get improvements. I mean, one thing that was kind of interesting is, so after we released this paper, there was another group of people, they actually just tried it on these open source models, and they actually also noticed the improvement, but it was like 2 % or 3 % compared to like our 20%. So I think there's like a pretty big disparity between like these commercial models and the ones that are available at the moment.
40:08Oh, that's fascinating. Yeah. And that's another topic that I'm really interested in is whether open source for generative AI is the future or whether the open source community just can't marshal the resources, the financial resources necessary to compete with proprietary models. And, you know, Meta is the only one that's sort of stepped forward to commit serious resources behind an open source LLM. But they're going to have to keep that funding at a very high level in order to compete. And do you think that open source generative AI will eventually, you know, that Yen Le Koon says that that's the future, that all these things will be open source.
41:08But then I talked to other people who say no, because the resources are concentrated in these very large companies and open source will never be able to compete. Meta will not be able to keep up with Anthropic or OpenAI by improving the open source models. What's your view on that? Yeah, I feel... Okay, so I do think at the moment, there's a pretty massive gap between the open source language models and the internal ones. I think a huge reason for this gap is actually because of data. So I think OpenAI has spent a lot of time gathering very customized data to make the agent look very appealing to people.
42:07And I think this lack of data, it feels like it's possible. I don't know. So I think one thing is that OpenAI and Anthropic, which I guess was from OpenAI actually, they were doing this large language model scaling for four or five years. And they were also looking for how to commercialize it for the last 2.5 years also. And I think everyone else kind of just suddenly realized they had to do this maybe a couple of months, I guess now at this point, almost nine months ago. But around nine months ago, people were suddenly like, oh, we need to spend time getting these large language models. So I think there's a big time difference at the moment.
42:58And I think, yes, the open source models are much worse. I think even the company models, I think Facebook's model internally, my guess is, I don't know, I don't know any information about this, is probably much worse than the OpenAI ones. And I think a lot of it is just a time thing. So I would hope that maybe in a year or two, I think, or maybe even in one year, maybe things will catch up. But yeah, at the current moment, things are very different. I think one thing that people seem to constantly say is they feel like there's a lot of technical knowledge to train these models. And then that's one reason why OpenAI is so much better than the other ones.
43:37I actually don't think that's the case. I do believe that... I mean, I think there are a couple of things that people don't appreciate about training the models. I think one big one is finding hyperparameters on smaller models and having the exact same hyperparameter scale for larger ones but like besides the a few couple technical tricks i think the main thing that is missing right now is just the data uh and like and like yeah so so i think because it's just the data i feel like maybe in a year or two people should catch up i don't think there's a technical know-how that's missing yeah and the gpu constraints and the cost of the compute that you don't think that's uh a barrier to open source uh yeah i think that could be a barrier so yeah it is true that the GPU cost is quite large, like probably like several, like maybe, maybe even a hundred million, I guess, like maybe hundreds of millions, maybe even a billion was like huge models.
44:32Yeah. I mean, I feel like maybe we need some type of like a research center or something like this where like government funded things. So like, I guess like with particle accelerators and particle physics, you have the large Hadron Collider and that's like hundreds of millions of dollars or more. So if we had an equal size effort, I think you could totally do this. Yeah, but actually, I do feel like at the moment, we don't even have... I feel like there's a disconnect between what people are... Yeah, it does feel like there's a big disconnect between what people constantly say online right now in the open source community.
45:09Everyone's like, oh, you just need a small model. You just need to fine tune it a little and then will reach their performance in the actual performance like uh yeah i think we people maybe should focus a bit more on using more gpus uh yeah that could be a thing that would be that could be problematic yeah uh and then uh you you you're also working uh in uh in robotics is is Is that right? Mm-hmm. So how does this work relate to robotics? Yeah, so I've been, so I guess throughout my PhD, I've been very interested in the idea of artificial intelligence. And in particular, I think that like to have artificial intelligence, we really need to have it physically in the world.
45:55So we need an agent that can visually understand the world, that can like physically act in the world, have like intents in the world and like interact with other agents. So I really believe in this physical AI hypothesis. So yeah, and I see that robotics is basically the instantiation of having a physically intelligent AI agent. So to me, I guess this multi-agent debate stuff, it's one way to try to get AI, to get a more intelligent AI. I think this idea in general could be applied a lot in different ways. So right now, I guess it's just on language models, but you can imagine that you have different models in the robot processing system, right?
46:38So maybe you have a video generation model that generates possible actions you want to do, right? You have some type of action model that takes images and predicts actions that you want, right? And you can imagine each of these models talk with each other. Like the action model says, you gave me this image. This image is not possible. Please refine it a little, right? So you can imagine that you have a network of these individual modules that all are talking with each other. And then like when you have like this set of these decentralized models that like are all talking with each other, maybe that whole thing can be like some like distributed operating system to like construct an intelligent robot.
47:12But like I feel like this idea of like debate or like the self-improvement thing like could be like broadly applied across like multimodal models like on a robot also. Yeah. Yeah, it seems, you know, I've had Peter Abiel, I'm sure you know, on the podcast, and Sergey Levine and different people working in robotics. and uh it it it seems that that the barrier in robotics is not so much uh the brain ai but it's it's just the the mechanics the miniaturization of motors and and you know all of that stuff do you feel that or or or is it is that uh is the hardware going to catch up uh with the uh with the software with the ai yeah i i think so i agree that hardware for robots is like a lot worse than what people have um yeah we don't have tactile sensing uh yeah the arms are very They are very awkward to use.
48:25They are very constrained. But I think when you give a person a remote control of the robot, they can do a lot of tasks in the kitchen. Almost every task you can imagine, you can actually do by remotely controlling the robot. So yes, those are big issues, but it feels like we're also missing the fundamental intelligent person also. I, as a person, can do a lot more than I can autonomously code my robot to do at the moment. And I think that if we could do more autonomously with the robots, then people would be much more interested in robotics. I think then we would have the interest in hardware to get better hardware to train robots.
49:06I think the thing is there's no proof of concept of AI working in robotics still that makes it so that people aren't very interested in trying to get better hardware for the robots. That's interesting. Yeah. Okay. Well, we're almost up to an hour. I don't want to keep you over, but are there areas of this that you're working on that I haven't asked about that listeners might be interested in? Oh, I mean, I think there's one kind of interesting idea that I feel like so I feel like every everyone, and it's kind of related to this thing I was talking about, about this, like, a multi-agent debate is a way to improve these language models.
49:49I think right now, everyone is working on one very giant monolithic model that can do everything. It can sense, it can perceive and stuff like this. But I feel like our brain is actually very decentralized. We have areas for motor control, we have areas for vision, we have areas for memory. So I think it's cool to think about this idea of having not a single AI, but have a society of AIs in your mind. Like you have a bunch of different small agents are talking to each other. And then like, that's how you actually function. Like you have one thought process that's trying to remember things that you're doing.
50:24Another one that kind of perceives what you're doing. So, yeah, I think that's an interesting idea that I feel like maybe people would also find interesting to think about. Yeah. And where did you do your undergraduate? I actually did my undergraduate at MIT also. Oh, OK. And and where are you going to go after this? Are you going to stay in academia? Yeah, I'm planning to stay in academia at the moment. Yeah, and I'm curious why. I mean, these days, you don't have to leave academia to go into industry. I mean, all these guys, Yan La Koon, NYU F. Meta, or Jeff Hinton, I mean, he's retired now, but Google and University of Toronto.
51:09I mean, there's a lot of opportunity to do both. Yeah, so I think if you had asked me this maybe one or two years ago, I would have been a lot less certain because it feels like at the time, yeah, it felt like you could do very open-ended research both in industry as well as in academia. I feel like nowadays, because there's so much commercial success with these models, the research you can do in industry is actually very limited. So if you want to work on RL agents without any language model, you can't really do that. There's a huge push for you to work on RL for language models. Or if you say, I don't believe that a single large language model will solve everything, which I don't know.
51:54I've been playing around. So during the summer, I was working at Google. Earlier, I was working at Facebook. I generally have been playing a lot with these vision language models. They still seem very, very far away from actual intelligence with the vision side. But like, even if you like, so if you don't believe in this hypothesis, you still have to work on this. Because that's the Vogue theme. And people think that it's so promising commercially that like, that's the idea to work on. So I feel like nowadays, you have very little freedom in industry. And I think it's actually just been in the last year that this has happened.
52:29Hi, this episode is sponsored by Salonis, the global leader in process mining. AI has landed and enterprises are adapting, giving customers slick experiences and the technology to deliver. The road feels long, but you're closer than you think. You see, your business processes run through many systems, creating data at every step. Salonis reconstructs this data to generate process intelligence, a common business language. With process intelligence, AI knows how your business flows across every department, every system, and every process. With AI solutions powered by Solonis, enterprises get faster, more accurate insights, a new level of automation, and a step change in productivity, performance, and customer satisfaction.
53:23Process intelligence is the missing piece in the AI-enabled tech stack. Search Celonis, C-E-L-O-N-I-S, to find out more. This episode is sponsored by Crusoe Cloud. High-performance cloud computing or low environmental impact, Crusoe Cloud was built because the innovations of the future need both. Crusoe Cloud is a scalable, clean, high-performance cloud optimized for AI and HPC workloads and powered by wasted, stranded, or clean energy. Crusoe offers virtualized compute and storage solutions for a range of applications, including generative AI, computational biology, and rendering. Visit crusoecloud.com, that's C-R-U-S-O-E-C-L-O-U-D dot C-O-M, to see what climate-aligned computing can do for your business.
54:31That's it for this episode. I want to thank Ilun for his time. If you'd like to read a transcript of today's conversation, you can find one, as always, on our website, I on AI, that's E-Y-E hyphen O-N dot A-I. And remember, the singularity may not be near, But AI is already changing our worlds, so pay attention.
From the publisher
This episode is sponsored by Crusoe. Crusoe Cloud is a scalable, clean, high-performance cloud, optimized for AI and HPC workloads, and powered by wasted, stranded or clean energy. Crusoe offers virtualized compute and storage solutions for a range of applications - including generative AI, computational biology, and rendering.
Visit https://crusoecloud.com/ to see what climate-aligned computing can do for your business
This episode is sponsored by Celonis ,the global leader in process mining. AI has landed and enterprises are adapting. To give customers slick experiences and teams the technology to deliver. The road is long, but you're closer than you think. Your business processes run through systems. Creating data at every step. Celonis recontrusts this data to generate Process Intelligence. A common business language. So AI knows how your business flows. Across every department, every system and every process.
Go to https:/celonis.com/eyeonai/ to find out more.
Welcome to episode 147 of the Eye on AI podcast. In this episode, host Craig Smith sits down with Yilna Du, a final year PhD student at MIT EECS with a background in research at leading institutions like OpenAI, FAIR, and Google Deepmind.
Yilun's extensive expertise spans generative models, decision making, robot learning, and embodied agents, making him a valuable voice in the AI domain.
Our conversation kicks off with a brief on Yilun's academic journey, leading into a deep dive into Reinforcement Learning with AI feedback (RLHF) - its history, inception, and challenges. We then touch upon the effectiveness of RLHF, the intriguing concept of multi-agent debate, and the PAPES procedure.
Craig and Yilun further explore the vast realm of AI, debating the gaps between open-source and proprietary models, the need for more compute resources, and the future of robotics interlaced with AI. Yilun provides a glimpse into his vision of decentralized AI systems, contrasting the industry's commercial trajectory with academia.
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview, Celonis and Crusoe Ad
(04:06) Yilun's Academic Background
(05:52) Origin and Applications of RLHF
(12:16) ROHF and the Multi-Agent Debate Method
(17:32) AI Model Interaction without Human Intervention?
(20:41) Applicability and Inconsistency Detection
(28:43) The Future of AI Training
(45:26) Robotics and Decentralized AI Systems




