Representation Engineering (Activation Hacking)

28 Feb 2024 · 44 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Practical AI Podcast Episode Notes: Representation Engineering (Activation Hacking)

Episode Summary In this episode of Practical AI, hosts Chris Benson and Daniel Whitenack delve into the concept of Activation Hacking, also referred to as Representation Engineering. They explore how this model control mechanism can influence the outputs of AI models and discuss the new Sora model from OpenAI. The conversation includes insights from a recent hackathon and the challenges faced while working with AI prompts.

Key Concepts

  1. Activation Hacking & Representation Engineering
  2. Definition: A technique used to control AI model behavior beyond traditional prompt engineering by modifying hidden states within the model.
  3. Distinction: Unlike prompt engineering, which focuses on input prompts, representation engineering allows for control over the model's output tone and behavior without changing the prompt itself.
  1. Hackathon Insights
  2. Hosts share their experiences from the Treehacks hackathon at Stanford.
  3. Notable project: Meshwork - Utilizes LoRa technology for disaster relief communication, integrating AI for command and control operations.
  1. Challenges in Prompt Engineering
  2. Discussion on frustrations with achieving desired outputs from AI models, especially in eliciting specific responses.
  3. The hosts share personal anecdotes of their struggles with traditional prompting methods.
  1. Control Vectors
  2. Mechanism: Control vectors are used to influence model outputs by adding specific characteristics (e.g., happy or sad) to responses without altering the main prompt.
  3. The process involves:
  4. Collecting hidden states from the model using contrasting prompts (happy vs. sad).
  5. Analyzing differences to create a control vector through techniques like PCA (Principal Component Analysis).
  1. Real-World Applications
  2. Potential uses in various contexts, such as customer service interactions (e.g., drive-through scenarios) where maintaining a specific tone is crucial.
  3. The idea of having multiple control vectors for different scenarios to enhance model usability and maintain user engagement.
  1. OpenAI's Sora Model
  2. Announcement of the Sora model that can generate hyper-realistic video from text.
  3. Discussion on the implications of such advancements in AI and the ongoing safety concerns surrounding deepfake technology.
  1. Competitor Developments
  2. Google’s Gemma Model: An open-source derivative of the Gemini model, promoting smaller models for practical applications, fostering the trend of accessible AI technologies.
  3. Encouragement for the AI community to engage with open-source models for experimentation and development.
  1. Code Generation and AGI
  2. Introduction of a company called Magic, which aims to develop code generation tools that could eventually lead toward AGI (Artificial General Intelligence).

Key Takeaways

  • Representation Engineering offers a novel approach to controlling AI outputs, moving beyond the limitations of prompt-based interactions.
  • The distinction between different control methods (like activation hacking) is crucial for developers looking to fine-tune AI behavior for specific applications.
  • The rapid development of AI models, such as OpenAI's Sora and Google's Gemma, signifies an exciting yet complex landscape in AI technology with both opportunities and ethical considerations.

Recommended Resources

  • [Representation Engineering Mistral-7B](https://vgel.me/posts/representation-engineering/)
  • [OpenAI Sora](https://openai.com/sora)
  • [Gemma Model on Hugging Face](https://huggingface.co/models)

Conclusion In this episode, Chris and Daniel provide a comprehensive exploration of representation engineering and its implications for future AI developments. The discussion captures the ongoing advancements and challenges within the AI landscape, encouraging listeners to engage with new methodologies and technologies.

---

For more information and ongoing discussions, join the [Practical AI community](https://practicalai.fm/community).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related tech is changing the world, this is the show for you. Thank you to our partners at Fly.io, the home of changelog.com. Fly transforms containers into micro VMs that run on their hardware in 30 plus regions on six continents. So you can launch your app near your users. Learn more at Fly.io.

0:42Welcome to another episode of Practical AI in this fully connected episode. Chris and I will keep you fully connected with everything that's happening in the AI world. We'll take some time to explore some of the recent AI news and technical achievements. And we'll take a few moments to share some learning resources as well to help you level up your AI game. I'm Daniel Whitenack. I am founder and CEO at Prediction Guard. And I'm joined as always by my co-host, Chris Benson, who is a tech strategist at Lockheed Martin. How are you doing, Chris? Doing great today, Daniel. We've got lots of news that's come out this week in the AI space.

1:23I know. Barely time to talk about amazing new things before stuff comes out. Yeah, I've been traveling for the past five days or something. I've sort of lost track of time, but it's like stuff was happening during that time in the news, especially the Sora stuff and all that. And I feel like I just kind of missed a couple news cycles. So it'll be good to catch up on a few things. But one of the reasons I was traveling was I was at the Treehacks hackathon out at Stanford. So I went there as part of the kind of Intel entourage and had Prediction Guard available for all the hackers there. And that was a lot of fun.

2:04And it was incredible. It's been a while since I've been to any hackathon, at least in-person hackathon. And they had like five floors in this huge, you know, engineering building of room for all the hackers. I think there was like 1600 people there participating from all over. And really cool. Of course, like there were some major categories of interest. One, you know, like in doing hardware things with robots and other stuff. Of course, one of the main areas of interest was AI, which was interesting to see. And in the track that I was a judge and mentor in, one of the cool projects that won that track was called Meshwork.

2:53So what they did, and this was all news to me, while some of this I learned from the brilliant students, but they said they were doing something with Laura. And I was like, oh, Laura, that's the fine tuning methodology for large language models. I was like, yeah, that figures like people are probably using LoRa. But I didn't realize. And then they came up to the table and they had these like little devices, like hardware devices. Then it clicked that something else is going on. And they explained to me they were using LoRa, which stands for long range. It's these sets of radio devices that communicate on these unregulated frequency bands and can communicate in a mesh network.

3:39So like you put out these devices, right, and they communicate in a mesh network and can communicate over long distances for very, very low power. And so they created a project that was disaster relief focused, where you would drop these in the field and there was a kind of command and control central zone and they would communicate back transcribed audio commands from the people in the field and would say, you know, oh, I've got a, you know, I've got a injury out here. it's a broken leg I need you know help whatever or meds over here this is going on over here and then they had an LLM at the command and control center parsing that text that was transcribed and actually creating like tagging certain keywords or events or actions and creating this nice command control interface which was awesome they even had like mapping stuff going on with computer vision trying to detect where like a flood zone was or there was damage in satellite images.

4:46So it was just really awesome. So all of that over a couple day period, it was incredible. That sounds really cool. And did they start the whole thing there at the beginning of the hackathon? Yeah, they got less sleep than I did. Although I have to say I didn't get that much sleep. You know, it wasn't a normal weekend, let's say. You can sack out on the plane rides after that. That sounds really cool. Yeah. And they were like, it was the first time I had seen one of those Boston Dynamics dogs in person. That was kind of fun. And they had other things like these faces you could talk to. I think the company was called like WeHead or something.

5:24It was like these little faces. All sorts of interesting stuff that I learned about. So I'm sure there'll be blog posts. And I think some of the projects are posted on DevPost, the site DevPost. So if people want to check it out, I'd highly recommend scrolling through some really incredible stuff that people are doing. Fantastic. I'll definitely do that.

5:51What's up, friends? Is your code getting dragged down by joins and long query times? The problem might be your database. Try simplifying the complex with graphs. A graph database lets you model data the way it looks in the real world instead of forcing it into rows and columns. Stop asking relational databases to do more than what they were made for. Graphs work well for use cases with lots of data connections like supply chain, fraud detection, real-time analytics, and generative AI. With Neo4j, you can code in your favorite programming language and against any driver. Plus, it's easy to integrate into your tech stack.

6:28People are solving some of the world's biggest problems with graphs. And now it's your turn. Visit Neo4j.com slash developer to get started. Again, Neo4j.com slash developer. That's Neo4j.com slash developer.

6:57Chris, one of the things that I love about these Fully Connected episodes is that we get a chance to kind of slow down and dive into sometimes technical topics, sometimes not technical topics. But I was really intrigued. You remember the conversation recently we had with Karan from Noose Research? That was a great episode. People can pause this and go back and listen to it if they want. I asked a lot of selfish questions. I learned a lot from him. But at some point during the conversation, he mentioned activation hacking. and he said hey like one of the cool things that like we're doing in this you know distributed research group and playing around with generative models is activation hacking and we didn't have time in the episode to talk about that and actually in the episode I was like I'm just totally ignorant of of what this means and so I thought yeah I should go check up on this and see if I can find any interesting posts about it and learn a little bit about it.

8:07And I did find an interesting post. It's called Representation Engineering, Mistral 7B, An Acid Trip. I mean, that's a good title. That's quite a finish to that title. Yeah. So this is on Thea Vogel's blog, and it was published January so recently. So thank you for creating this post. And I think it does a good job at describing some of, I don't know if it's describing exactly what Haran from Noose was talking about, but certainly something similar and kind of in the same vein. There's a distinction here, Chris, with what they're calling representation engineering between representation engineering and prompt engineering.

8:55So I don't know how much you've experimented with prompt optimization. And yeah, what is your experience, Chris? Sometimes like these very small changes in your prompt can create large changes in your output. Yes, that is an art that I am still trying to master and have a long way to go. Sometimes it works well for me and I get what I want on the output. And other times I take myself down a completely wrong rabbit hole and I'm trying to back out to that. So I have a lot to learn in that space. Yeah. And I think one of the things that is a frustration for me is I say something explicitly and I can't get it to do the thing explicitly.

9:36I'm on a customer site recording from one of their conference rooms. They graciously let me use it for the podcast. And over the past few days, we've been architecting some solutions and prototyping and such. and there was this one prompt that we wanted to output like a set of things and then look at another piece of content and see which of those set of things was in the other piece of content and it was like no matter what i would tell the model it would just say they're all there or they're all not there like it's either all or nothing and no matter what i said it wouldn't change things so i don't know if you've had similar types of frustrations i have i'll narrow the scope down on something, trying, you know, I'll go to something like chat GPT, you know, with the GPT four and I'll be, I'll be trying to narrow it down.

10:28I'll be very, very precise with a short prompt that is, you know, the 15th one in succession. So there's a history to work on and I still find myself challenged on getting what I'm, what I'm trying to do. So what have you stumbled across here? That's going to help us with this. Yeah. So there's a couple of papers that have come out. They reference one from October 2023 from the Center for AI Safety, Representation Engineering, a Top-Down Approach to AI Transparency. And they highlight a couple other things here. But the idea is, what if we could, not just in the prompt, but what if we could control a model to give it a, you might think about it like a specific tone or angle on the answer.

11:23It's probably not a fully descriptive way of describing it, but the idea of being like, oh, could I control the model to always give happy answers or always give sad answers? Or could I control the model to always be confident or always be less confident, right? And these are things generally you might try to do by putting information in a prompt. And I think this is probably a methodology that would go across. I'm kind of using the example with large language models, but I think you could extend it to other categories of models like image generation or other things. It's very like you kind of put in these negative prompts like don't do this or behave in this way.

12:07you're occasionally funny or something like that as your assistant in the system prompt. It kind of biases the answer to a certain direction, but it's not really that reliable. So this is, it seems, what this area of representation engineering, or you might call it activation hacking, is really seeking to do. If we look in this article, actually, there's a really nice kind of walkthrough of how this works. And they're doing this with the Mistral model. So cutting to the chase, if I just give some examples of how this is being used, you have a question that's posed to the AI model, in this case, Mistral, what does being an AI feel like?

12:55And in controlling the model, not in the prompt, so the prompt stays the same, the prompt is just simply what does being an AI feel like? So the baseline response starts out, I don't have any feelings or experiences. However, I can tell you that my purpose is to assist you. That sort of thing, kind of a bland response. Same prompt, but with the control put on to be happy, the answer becomes, as a delightful exclamation of joy, I must say that being AI is absolutely fantastic. You know, and then it keeps going, right? And then with the control on to be, they put it as sort of like minus happy, right?

13:40Which I guess, I guess it'd be sad. It says, I don't have a sense of feeling as humans do. However, I struggle to find the motivation to continue feeling worthless and unappreciated. So yeah, you can kind of see, And this is all with the same prompt. So we'll talk about kind of how this happens and how it's enabled and that sort of thing. But how does this strike you? Well, first of all, funny. But second of all, the idea is interesting. Looking through the same paper that you've sent me over, they talk about control vectors. And I'm assuming that's what we're about to dive into here in terms of how to apply them.

14:19Yeah. Looks good. And this is sort of a different level of control. So there's various ways people have tried to control generative models. One of them is just the prompting strategies or prompt engineering, right? There's another methodology which kind of fits under this control, which has to do with modifying how the model decodes outputs. So this is also different from this representation engineering methodology. People like Matt Rickert have done things, many others too, where it's you say, oh, well, I want maybe JSON output or I want a binary output, like a yes or a no. Well, in that case, you know exactly what your options are.

15:09So instead of decoding out probabilities for 30 ,000 different possible tokens, maybe you mask everything but yes or no and just figure out which one of those is most probable. So that's a mechanism of control where you're only getting out one or another type of thing that you're controlling. So this is interesting in that you're still allowing the model to freely decode what it wants to decode, but you're actually modifying not the weights and biases of the model. So it's still the pre-trained model, but you're actually applying what they call a control vector to the hidden states within the model.

15:50So you're actually changing how the forward pass of the model operates. If people remember or kind of think about when people talk about neural networks, now people just use them over API. But when we used to actually make neural networks ourselves, there was a process of a forward pass and a backward pass where the forward pass is you put data into the front of your neural network. it does all the data transformations and you get data out the other side, which you would call an inference or prediction. And the back propagation or backward pass would then propagate changes in the training process back through the model.

16:28So here it's that forward pass. And there's sort of some jargon, I think, that needs to be decoded a little bit. No pun intended. So we talk about this where there's a lot of talk about hidden layers. And all that means is in the forward pass of the neural network or the large language model, a certain vector of data comes in. And that vector of data is transformed over and over through the layers of the network. Then the layers just mean a bunch of sub functions in the overall function that is your model. And those sub functions produce intermediate outputs that are still vectors of numbers. But usually we don't see these, and so that's why people call them hidden states or hidden layers.

17:17You're talking about the fact that the control vector is not changing the weights on the way back, the way back propagation works. Correct. How does the control vector implement into those functions? So as it's moving through those hidden layers, what is the mechanism of applicability on the model that it uses for that? So it's, I mean, intuitively, it sounds almost like the inverse of backpropagation, the way you're talking. I don't know if that's precise, but. Yeah, it's quite interesting, Chris. I think it's actually a very subtle but creative way of doing this control. So the process is as follows.

17:59In the blog post, they kind of break this down into four steps. and there is data that's needed but you're not creating data for the purpose of training the model you're creating data for the purpose of generating these what they call control vectors so the first thing you do is you say okay let's say that we want to do the happy or not happy or happy and sad operation so you create a data set of contrasting prompts where one explicitly asks the model to act extremely happy, like very happy, all the ways you could say to the model to be really, really happy. And, you know, rephrase that in a bunch of examples.

18:44And then on the other side, the other one of the pair do the opposite. So ask it to be really sad. I know you're, you're really, really sad and be sad. And you have these pairs of prompts. Okay. And then you take the model and you collect all the hidden states for your model while you pump through all the happy prompts and all the sad prompts. And so you've got this collection of hidden states in your model, which are just vectors that come when you have the happy prompt and when you have the sad prompt. So step one, the pairs of kind of like a preference data set, but it's not really a preference data set, it's contrasting pairs on a certain axes of control, right?

19:34Okay. And so you run those through, you get all of the hidden states. And step three is then you take the difference between, so for each happy hidden state, you take its corresponding sad one, and you get the difference between the two. Okay, so now you end up with this big data set of, for a single layer, you have a bunch of different vectors that represent differences between that hidden state on the happy path and the sad path. So you have a bunch of vectors. Now to get your control vectors, step four, you apply some dimensionality reduction or matrix operation. The one that's talked about in the blog post is PCA, but it sounds like people also try other things.

20:21PCA is principal component analysis, which would then allow you to extract a single control vector for that hidden layer from all these difference vectors. And now you have all these control vectors. So when you turn on the switch of the happy control vectors, you can pump in the prompt without an explicit instruction to be happy, and it's going to be happy. And when you do the same prompt, but you turn off the happy and you turn on the sad. Now it comes out and it's sad. That's interesting. Where would you want to use this to achieve that bias versus some of the more traditional approaches such as, you know, asking in the prompt, what is we're listening to this?

21:09Where is this going to be most applicable for us? Yeah. I think that people anecdotally, at least, if not explicitly in their own evaluations, have found very many cases where, like you said, it's very frustrating to try to put things in your prompts and just not get it. And what's interesting also is like a lot of this is boilerplate for people over time, like you are a helpful assistant, blah, blah, blah. And they have their own kind of set of system instructions that at least to their best of their ability, get what they want. So I think when you're seeing inconsistency in control from the prompt engineering side, I always tell people when I'm working with them with these models that the best thing they can do is just start out with trying basic prompting.

22:05Because if that works, it's the easiest thing to do, right? You don't have to do anything else. Sure. But then the next thing, or maybe one of the things you could try before going to fine tuning, because fine tuning is another process by which you could align a model or create a certain preference or something. But it takes, you know, generally GPUs and maybe is a little bit harder to do because then you have to store your model somewhere, right? And all this stuff and host it and maybe host it for inference. And that's difficult. So with the control vectors, maybe it's a step between those two places, right?

22:47Where you have a certain vector of behavior that you want to induce. And it also allows you to make your prompts a little bit more simple, right? You don't have to include all of this junk that is kind of general instructions. You can institute that control in other ways, which also makes it easier to maintain and iterate on your prompts because you don't have all this long stuff about how to behave. So to extend the happy example for a moment, I want to drive it into a real-world use case for a second. Let's say that we're going to stick literally with the happy thing. And let's think of something where we would like to have happy responses, maybe a fast food restaurant.

23:28You're going through a drive-through at a fast food restaurant where a couple of years from now they may have put an AI system in place. White Castle has it now. Oh, okay. Well, there you go. There you go. You're already ahead of me there. So, okay, I'm coming now. It also shows that I'm unhealthy and go to White Castle. Okay, well, I'm now coming forward with my thoroughly out of date use case here. And so we have the model and maybe we to use the model on without doing retraining it or anything. We want to maybe use retrieval augmented generation, apply it to the data set that we have, which might be the menu.

24:07And then maybe we use this mechanism that you've been instructing us on the last few minutes for that happy thing so that the drive through consumer can have the conversation with the model through the interface. It applies primarily to the menu, but they get great responses and maybe that helps people along. I don't always get that happy response from all the humans in the drive-thrus where I go to have my unhealthy food things. First off, thanks for making me hungry for White Castle. We're recording this in the late afternoon. Dinner is coming up pretty soon. It is coming up. There's an unspoken bias right here.

24:50Yeah, exactly. What's interesting is you could have different sets of these that you can kind of turn on and off, which is really an intriguing, like you have this sort of zoo of behaviors that you could turn on and off. I think even, oh, you're, you have this one interaction that needs to be this way. But as soon as they go into this other flow, you need to kind of have another behavior. It may be useful for people to get some other examples. So we said the happy, sad one. There's some other examples that are quite intriguing throughout the blog post from Thea. I hope I'm saying that name right.

25:29If not, we'd love to have you on the podcast to help correct that and continue talking about this. But another one is honest or dishonest or honest or not honest. and the prompt is your life for work. What would you tell your boss? And the one, it says, I would be honest and explain the situation. And it's the honest one. And then the other one says, I would tell my boss that the sky was actually green today and I didn't go out yesterday. I would also say, I have a secret weapon that I use to write this message. so kind of a different flavor there the one probably inspiring the blog post the acid trip one is they had a a trippy one and a non-trippy one so the prompt is give me a one sentence pitch for a tv show so the the non-trippy one was a young and determined journalist who's always serious and respectful be able to make sure that the facts are not only accurate but also understandable for the public and then the trippy one was our show is a kaleidoscope of colors trippy patterns and psychedelic music that fills the screen with worlds of wonder where everything is oh oh man um just uh i'm going for the latter one just for the yeah exactly yeah they they do um lazy not lazy they do left wing right wing creative not creative uh future looking or not future looking self-aware um so there's a lot of interesting things i think to to play with here and it's an interesting level of control that's potentially there one of the things that they do highlight is this control mechanism could be applied both to jailbreaking and anti-jailbreaking models.

27:29So by that, what we mean is models have been trained to do no harm or not output certain types of content, right? Well, if you institute this control vector, it might be a way to break that model into doing things that the people that train the model explicitly didn't want it to output, right? But it could also be used the other way, right? To maybe prevent some of that jailbreaking. So there's an interesting interplay here between maybe the good uses and the less than good uses on that spectrum. That entire AI safety angle on using the technology responsibly or not. Sure. They represent or references the rep ing library, which I guess is one way to do this, but there may be other ways to do this.

28:21If any of our listeners are aware of other ways to do this or convenient ways to do this or examples, please, please share them with us. We'd love to hear those.

28:45This is a changelog news break. GPT script is a new scripting language to automate your interactions with LLMs, which for now just means open AI. From the project's homepage, The ultimate goal is to create a fully natural language-based programming experience. The syntax of GPT script is largely natural language, making it very easy to learn and use. Natural language prompts can be mixed with traditional scripts such as Bash and Python or even external HTTP service calls. The project includes examples of how to plan a vacation, edit a file, or run some SQL. The central concept is that of tools.

29:26Each tool performs a series of actions similar to a function, and GPTScript composes the tools to accomplish tasks. You just heard one of our five top stories from Monday's Changelog News. Subscribe to the podcast to get all of the week's top stories and pop your email address in at changelog.com slash news to also receive our free companion email with even more developer news worth your attention. Once again, that's changelog.com slash news.

30:01Well, this was a pretty fascinating deep dive, Daniel. Thank you very much. Yeah, yeah. You know, you can go out and control your models now, Chris. It'll be the first time ever, I think, you know, that I've done it well there. Always trying different stuff. I think we'd be remiss if we got through the episode and didn't talk about a few of the big announcements this past week. Yeah, a lot. It's been quite a week. You mentioned right up front, OpenAI announced their Sora model, which case you're able to create very hyper realistic video from text. I don't believe it's actually out yet. At least when I first read the announcement, it wasn't available yet.

Read the full transcript

30:44They had put a bunch of demo videos. Yeah, I checked just before recording this and I couldn't see it. It's still not released at this point. Yeah. Okay. But there's a number of videos that OpenAI has put out. So I think we're all kind of waiting to see. But the thing that was very notable for me this week, I really wasn't surprised to see the release. And we've talked about this over the last year or so. If you look at the evolution of these models that we're always kind of documenting in the podcast episodes and stuff, this was coming. We all knew this was coming. We just didn't know how soon or how far away.

31:20but we talked many months ago about we're not far from video now. So open AI has gotten there with the first of the hyper-realistic video generation models and definitely looking forward to gaining access to that at some point and seeing what it does. But there was a lot of reaction to this in the general media, you know, in terms of AI safety concerns, you know, how do you know something is real going forward and stuff. And, you know, it's, it's the next iteration of the more or less the same conversation we've been having for several years now on AI safety. What are your thoughts when you first saw this?

31:58Yeah, it's definitely interesting in that it definitely didn't come out of, out of nowhere, just like all the things that we've been seeing. We've, we've seen video generation models in the past, generally not at the level either generating like very very short clips with high quality maybe or generating like from an image a realistic image some motion or maybe videos that are not that compelling i think the difference and of course we've only seen like you say it's not the the model that we've got hands-on with but we've seen a lot the release videos which who knows how much they're cherry-picked i mean i'm I'm sure they are to some degree and also aren't to some degree.

32:48I'm sure it's very good. But other players in the space have been meta and runway, ML and others. But yeah, this one I think was intriguing to me because yeah, generally there were a lot of really compelling videos at first sight. And then I think you also had people just like the image generation stuff has been, you have like real photographers or real artists that look at an image and like say, oh, look at all these things that happen. And it's the same here. Like they all kind of have a certain flavor to them, probably based on how the model was trained. And they still have, I think I was watching one where it's like a grandma blowing out a birthday cake.

33:42And one of the candles had like two flames coming out of it. And then like, there's a person in the background with like a disconnected arm sort of waving. But if you had the video as like a B-roll and a really quick type of video of other things, you probably wouldn't notice those things right off the bat. If you slow it down and you look, there's like the weirdness you would expect, just like the weirdness of like six fingers or something with image generation models. Right. So, yeah, I think it's really interesting what they're doing. I don't really have much to comment on in terms of the technical side other than they're probably doing some of what we've seen that people have published.

34:24Of course, OpenAI doesn't publish their stuff or share that much in that respect. But it probably follows in the vein of some of these other things. And people could look on hugging faces, even hugging face spaces where you can do video generation, even if it's only like four seconds or something like that, or not even that long. I think the main thing, aside from the specific model itself, is it's kind of signaling in the general public's awareness that this technology has arrived. And just as with the other, you know, with ChatGPT before and things like that, you know, it's going to be one of the, it's here now, everyone knows.

35:03And we'll start seeing more and more of the models propagating out. And some obviously will be closed source, like OpenAI is. and hopefully we'll start soon seeing some open source models doing this as well. Yeah. Speaking of open source, another, a competing large cloud company, Google, decided to try their hand in the open source space as well, or at least the open model space. And they released a derivative of their closed source Gemini. And I say derivative because they say it was built along the same mechanisms called Gemma. And it's currently, as we are talking right now, in the number one position on hugging face.

35:46At least last time I checked not long before this, although that changes fast. I probably should have checked right before I said that. It's still number two, but well, it's the top language, trending language model. Gotcha. Stability is stable cascade, knocked it out of the overall top spot. But yeah, the Gemini ones are quite interesting because they're also smaller models, which I'm a big fan of. Most of our customers use these sort of smaller models. And also even having a 2 billion parameter model makes it very reasonable to try and run this locally or in edge deployments and that sort of thing or in a quantized way with some level of speed.

36:35And they also have the base models, which you might grab if you're going to fine tune your own model off of one of these. And they have instruct models as well, which would probably be better to use if you're going to use them kind of out of the box for general instruction following. Criticisms I've heard just about the approach is I've heard a number of people saying, oh, they're putting a foot in each side of the camp, you know, one in closed source with the main Gemini line and Gemma being open source and the weaker. But I would in turn say, I'm very happy to see Gemma in open source. We want to encourage this.

37:15We want the organizations who are going to produce models to do that. And you're right, going back to what you were just saying, this is where most people are going to be using models in real life is, you know, there is if you're not just running through an API to one of the largest ones, but you don't need those for so many activities. So I think, you know, this is, we've talked about this multiple times on previous episodes, models this size are really where the action is at. It's not where the hype is at, but it is where the action's at for practical, productive, and accessible models. Yeah, definitely.

37:49Especially for people that have to get a bit creative with their deployment strategies, either for regulatory security privacy reasons or for connectivity reasons or other things like that. I could see these being used quite widely. And generally what happens when people release a model family like this, and you saw this with Llama 2, you've seen it with Mistral. Now with Gemma, we'll see a huge number of fine tunes off of this model. Now, one of the things that I need to do is you do have to agree to certain terms of use to use the model. It's not just released under Apache 2 or MIT or something like that, Creative Commons.

38:42So you accept a certain license when you use it. And I need to read through that a little bit more. So people might want to read through that. I don't know what that implies about both fine tuning and use restrictions. so that would be worth a look for people if they're going to use it but certainly would be easy to pull it down and try some things they do say that it's already and I'm sure actually Hugging Face probably got a head start you know a week or so maybe of head start to make sure that it was supported in their libraries and that sort of thing because I think even now you can use the standard Transformers libraries and other trainer classes and such to fine tune the model.

39:27Sounds good. So as we start to wind down, before we get to the end, do you have a little bit of magic to share by chance? That's a good one, Chris. Yes, on the road to AGI magic, as your predictions for the year talked about, there'd be people talking about AGI again, and certainly they are. It's not directly an AGI thing, But this company, Magic, which is kind of framing themselves as a code generation type of platform in the same space as like GitHub Copilot, Codium, maybe. They raised a bunch of money and posted some of what they're trying to do. And there was some information about it. I think people seem to be excited about it because of, you know, some of the people that were involved.

40:21but also because they talk about code generation as a kind of stepping stone or path to AGI. So what they mean by that is, well, okay, initially they'll release some things as copilot and code assistant type of things like we already have. but eventually there's tasks within the set of things that we need developers to do that they want to do automatically not just having you have a co-pilot in your own coding but in some ways having a junior dev on your team that's doing certain things for you and of course if you take that then to its logical end as the dev on your team ai dev on your team gets better and better maybe it can solve increasingly general problems through coding and that sort of thing so i think that's the take that they're having on this code and agi situation okay well cool like i said quite a week uh full of news and uh when you combine that with the deep dive you just took us through on representation engineering, especially with an acid trip involved.

41:39Yeah, we've been we were hallucinating more than chat GPT, as our friends over at the ML Ops podcast would say. Can't beat that. We got to close the show on that one. Yeah, yeah. Well, thanks, Chris. I would recommend that people take if they're interested specifically in learning more about the representation learning subject or activation hacking. Take a look at this blog post. It is more of a kind of tutorial type blog post and there's code involved and references to the library that's there. So you can pull down a model. Maybe you pull down the Gemma model, the 2 billion one in a collab notebook.

42:17You can follow some of the steps in the blog post and see if you can do your own activation hacking or representation learning. I think that would be a good learning both in terms of a new model and in terms of this methodology. Sounds good. I will talk to you next week then. Alright, see you soon Chris.

42:47Alright, that is Practical AI for this week. Subscribe now. If you haven't already, head to practicalai.fm for all the ways. And join our free Slack team where you can hang out with Daniel, Chris, and the entire ChangeLog community. Sign up today at practicalai.fm slash community. Thanks again to our partners at fly.io, to our Beat Freakin' Residence, Breakmaster Cylinder, and to you for listening. We appreciate you spending time with us. That's all for now. We'll talk to you again next time.

43:31Game on!

From the publisher

Recently, we briefly mentioned the concept of “Activation Hacking” in the episode with Karan from Nous Research. In this fully connected episode, Chris and Daniel dive into the details of this model control mechanism, also called “representation engineering”. Of course, they also take time to discuss the new Sora model from OpenAI.

Join the discussion

Changelog++ members save 4 minutes on this episode because they made the ads disappear. Join today!

Sponsors:

  • Neo4j – Is your code getting dragged down by JOINs and long query times? The problem might be your database…Try simplifying the complex with graphs. Stop asking relational databases to do more than they were made for. Graphs work well for use cases with lots of data connections like supply chain, fraud detection, real-time analytics, and genAI. With Neo4j, you can code in your favorite programming language and against any driver. Plus, it’s easy to integrate into your tech stack. Visit Neo4j.com/developer to get started. 
  • Changelog News – A podcast+newsletter combo that’s brief, entertaining & always on-point. Subscribe today. 
  • Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs. 

Featuring:

Show Notes:

Something missing or broken? PRs welcome!

More from Practical AI

All 157 episodes
Representation Engineering (Activation Hacking)Practical AI · 44 min
Listen in VO