There's a new Llama in town

25 Jul 2023 · 48 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Practical AI Episode Notes

Episode Title

There's a New Llama in Town

Podcast Description Practical AI focuses on making artificial intelligence practical, productive, and accessible to everyone. The show features discussions among technology professionals, business people, students, enthusiasts, and expert guests on AI-related topics, including Machine Learning, Deep Learning, Neural Networks, GANs, and more.

Episode Summary In this episode, hosts Chris Benson and Daniel Whitenack discuss significant developments in AI, including the release of Meta's Llama 2 and Google's new Zip-NeRF technology. They provide a comparative analysis of recent advancements in AI functionalities, including OpenAI's offerings and Anthropic’s Claude 2.

---

Key Topics Discussed

  1. Zip-NeRF Technology
  2. Overview: Zip-NeRF (Neural Radiance Field) generates 3D scenes from 2D images, enabling impressive visualizations.
  3. Use Cases:
  4. Real estate (creating 3D walkthroughs from 2D images).
  5. E-commerce (allowing customers to visualize furniture in their homes).
  6. Industrial training simulations.
  7. Technological Impact: Represents a shift in generative AI beyond text, showcasing significant 3D scene creation capabilities.
  1. Llama 2 Release
  2. Background: Meta's Llama 2 is a large language model (LLM) with a commercial license, allowing broader use compared to its predecessor.
  3. Features:
  4. Available in three sizes: 7 billion, 13 billion, and 70 billion parameters.
  5. Enables users with fewer than 700 million monthly active users to utilize it commercially.
  6. Technical Insights:
  7. Uses separate reward models for helpfulness and safety during fine-tuning.
  8. Restrictive licensing may limit how outputs can be used in developing other models.
  1. Comparison with Anthropic’s Claude 2
  2. Claude 2 Features:
  3. Larger input size capability (up to 100,000 tokens).
  4. Allows users to upload files directly for integrated context in responses.
  5. Methodology: Claude 2 generates outputs differently than other models by using the uploaded data within its prompts, contrasting with OpenAI's code interpreter, which generates and executes code based on user inputs.
  1. OpenAI's Code Interpreter
  2. Functionality: Processes uploaded data by generating code (e.g., Python) and executing it to return results.
  3. Contrast with Claude 2: OpenAI's approach is more about executing code for data analysis, while Claude 2 directly incorporates uploaded data into its processing.
  1. General Observations on AI Developments
  2. Rapid advancement in AI technologies demands a continuous evaluation of tools and methodologies.
  3. Importance of hands-on experimentation with different models to determine their suitability for specific use cases.
  4. The potential for both large organizations and small businesses to leverage these technologies for innovative applications.

---

Key Takeaways

  • Technological Evolution: The AI landscape is evolving rapidly, with significant advancements in generative models affecting industries from real estate to e-commerce.
  • Accessibility and Commercial Licensing: Llama 2's commercial license structure opens new opportunities for businesses at various scales, encouraging innovation without the high costs typically associated with LLMs.
  • Diverse Use Cases: Both Zip-NeRF and Llama 2 have wide-ranging applications that can benefit various sectors, indicating a shift towards more practical AI use cases.

Learning Resources

  • [What is NeRF?](https://datagen.tech/guides/synthetic-data/neural-radiance-field-nerf/)
  • [Llama 2 Overview](https://ai.meta.com/llama/)
  • [OpenAI Code Interpreter](https://openai.com/blog/chatgpt-plugins#code-interpreter)
  • [Claude 2 Overview](https://www.anthropic.com/index/claude-2)
  • [Hugging Face's Guide to Llama 2](https://huggingface.co/blog/llama2)
  • [Ethan Mollick's Blog on Code Interpreter](https://www.oneusefulthing.org/p/what-ai-can-do-with-a-toolbox-getting)

---

Conclusion This episode of Practical AI highlights the transformative developments in AI, focusing on new technologies that enhance productivity and real-world applications. The discussion emphasizes the importance of understanding diverse AI models and leveraging their unique features for various business needs.

Call To Action Listeners are encouraged to experiment with Llama 2, Zip-NeRF, and Claude 2, and to explore the provided resources for deeper insights into these technologies.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:06Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related technologies are changing the world, this is the show for you. Thank you to our partners at Fastly for shipping all of our pods super fast to wherever you listen. Check them out at Fastly.com. And to our friends at Fly, deploy your app servers and database close to your users. No ops required. Learn more at fly.io.

0:43Welcome to another fully connected episode of Practical AI. In these episodes, Chris and I keep you fully connected with everything that's happening in the AI community. We're going to take some time to discuss the latest AI news, and then we'll share some learning resources to help you level up your machine learning game. This is Daniel Whitenack. I'm a founder and data scientist at Prediction Guard, And I'm joined, as always, by my co-host, Chris Benson, who is a tech strategist at Lockheed Martin. How are you doing, Chris? Doing cool. I'm trying to figure out how did we survive before all these great new models and stuff?

1:22Like, it's changed. Yeah, it's been crazy. Yeah, I just created a post for LinkedIn and I was like grabbing text, putting it into chat GPT, like getting nice rephrasing and then I'm like oh I need an image and in particular we'll talk about it a little bit in this episode but I was like oh there's this free willy model from stable AI which is like whale themed right and then I've got the llama thing so I just went to stable diffusion excel on clip drop and said hey generate me an image with a whale and a llama together and How do I even post to LinkedIn before without these things? It's like a different world.

2:07Yeah. 2023 versus 2022 is totally different. The content generation, the way you code, it's a different world. Yeah. And this week, as most weeks are, it seems like in 2023, had some pretty groundbreaking announcements and releases, which we're going to dive into a bunch of those things. There's just a huge amount to update on. And I think it's a good time for one of these episodes between you and I to just parse through some of the new stuff that is hitting our feeds. So, yeah. Well, I mentioned Llama. One of the big things this week was Llama 2. But I think before we jump into Llama 2, which I think was maybe the main thing dominating at least my world this week, it might be worth just taking a little bit of time to highlight something outside of the stream of large language models, which also crossed my desk this week, which I thought was really cool.

3:09It's this latest version of Nerf. This is work from Google presented at ICCV 2023. So it's called ZIPNERF, Anti-Alias Grid-Based Neural Radiance Field. That's quite a name right there. It is quite a name. It stands for Neural Radiance Field. So NERF, it's like camel-cased in capital N, small E, and capital RF. NERF, these are fully connected neural networks that create unique novel views of complicated 3D scenes based on a set of images that are input. So I don't know if you've seen that video yet. I'm looking at it as we are talking. And when you say the video, I know which video you're talking about because it's amazing.

4:02I've just left it on. It's pretty spectacular. Yeah. So this is a podcast, so it's hard to express some of this for people. If you just search for zip nerf you can go to the page for this paper which is a great summary but there's a video on the page and just to describe what it is imagine this kind of complicated house with a bunch of different rooms and an outdoor patio sort of garden area and the video is actually this kind of almost like a drone fly through of the house and then the outdoor area if you imagine a drone flying through a house there's hats and coats and toys and couches and plants and all sorts of things everywhere but the video is extremely seamless and it's not generated by a drone it's actually just generated by interpolating between a whole bunch of 2d images and then interpolating from that the 3d scene so yeah i don't know what are your impressions chris first of all the the from the perspective, the drone flight, if you will, that you have as a perspective viewing it is like the best drone operator in the history of the world.

5:16Yeah. That'd probably be hard to get one to do that. Yeah. Yeah. You're not going to get a real drone operator that could fly that amazingly, you know, and get those things. It's just phenomenal. And the house is like, for a moment you look at it and I mean, it is, it looks real, but I have noticed it's cluttery, but it's immaculately clean at the same time as well. The clutter is cleanly distributed and stuff. So I wish when my house was cluttered, it looked as beautiful as this house. It doesn't. But yeah, I mean, just like if you didn't know, if you weren't listening to the Practical AI podcast to go look at it or something like that and you just stumbled upon it, you'd think it was a drone video if you didn't have the education and go, oh my God, this is just really cool.

6:02I wonder what they're doing here. But it's indistinguishable from real life for all practical purposes. Yeah, and so it's based on 2D images, and then there are these generated interpolations, which maybe gets to, there was something that we were talking about prior to hitting the record button, which was this whole field of generative AI is sometimes conflated with large language models or chat GPT, But there's a whole lot going on in generative AI that's not language related or maybe even based on language related prompts. So I mentioned that image that I generated for my LinkedIn post that was still a text prompt into a model that generated an image.

6:48But here what we're seeing is we've got static 2D images that are input to a model that's actually generating a whole bunch of different perspectives that are synthesized in a 3D scene. So this is, I would say, still fitting into our current landscape and world of generative AI. But it's not a text in, text out or text in, image out model. Right. And I think people, there's so much coming at people right now. I think, you know, we keep talking about that this year. In the five years we've been doing this podcast, we've never had a moment like the last few months where things have been coming, new things have been coming at people so fast, new terms, new models, and people are trying to distinguish.

7:36So it's pretty, I think it's pretty fair that people are trying to make sense of how they relate together. And there's, there's a lot of connecting between, you know, the, the idea of generative and the idea of large language models overlap in a lot of areas. And you have models that are both and you have models that are just one and stuff. But I think it's a brave new world right now in terms of the, the amount, every show, we're just trying to figure out what matters right now. Cause there's a lot we're not hitting. Yeah. And this side of things, maybe like the 3D or video or image based side of things, I know has its own set of kind of transformative use cases that are popping out.

8:11I even remember a little while ago there was some technology, I think from Shopify, but others have done this as well, where maybe you have a room in your house and you want to see how you can transform it with new furniture or something that, of course, you could buy. This is a real kind of e-commerce or retail sort of use case for some of the scene technology of a different kind. If you think of this sort of technology that can take 2D things and create these 3D scenes, certainly there's use cases within game development, for example. But even other cases where maybe AI has never impacted the process as much like in real estate, for example, you know, how expensive is it to literally have a person come out with specialized camera gear?

9:08I know that we've had this in the past where it takes a special person to come out with special camera gear to capture the kind of 3D walkthrough, essentially the street view walkthrough of your house and map that onto an actual schematic of your house. And here, if you imagine someone, maybe I'm now selling my house myself without a real estate agent, and I can take an app potentially and go through my house just taking 2D images and create this really cool kind of fly around 3D view that's interactive. That's really, I think, a powerful transformative change for a number of different industries.

9:51I came across a company called Luma AI in one of the posts about this technology. I don't know exactly how much of the, if they're even using the Zip Nerf stuff, but certainly some things related to Nerf to take these 2D images and they have an app that will create 3D views, which is pretty cool to see some of this kind of hit actual real users. We keep talking about the fact that we've hit this inflection point where it's hitting all the, you don't have to be in the AI world, you know, for this to have a big impact. And so, you know, it's very easy looking at the ZipNerf video to imagine walking around with your cell phone on an app and you're, you're just kind of like walking around and the app takes care of whether it's video or whether it's still images or what, and it just uploads it to this and produces this, you know, amazing, you know, so it's not your walk around that it's doing.

10:46It takes that as raw video, but then it produces this super high quality thing. So yeah, I mean, I think this is another case where there's this one technology with thousands of use case possibilities, you know, where it just changes everything. Yeah. And maybe also in the, it'd be curious to know your reaction to this also with respect to kind of the industrial use cases where I've been thinking about, of course, Just like capturing 3D scenes is very important, for example, for simulated environments where you're trying to maybe train an agent or you even kind of an industrial training for human sort of sort of scenario where you want to kind of take someone into an environment that it's physically hard to bring a lot of people into.

11:36Or there could be safety issues and such. Yeah, safety issues. I don't know if that sparks things in your mind. I think in the industrial sense, this could have a more B2B sort of impact than just a consumer app. Sure. I mean, a simple thing, and I'm making something up in the next thing I say, but it's very easy for me to imagine intelligence agencies that are, you know, like if you go back some years to when Osama bin Laden was found and they had various imagery and stuff. But with stuff like this, they might take all those images that they're getting from various sources and produce, you know, a high.

12:15Like a flyover. Yeah, a flyover and very photorealistic of certain parts of the compound where that kind of imagery and that can be used in a military operation subsequently. Now, I'm making that up. So nobody should take that as a thing. But it's not hard to imagine that. It's not hard to imagine a lot of factory uses and other industrial things where you have safety issues, you have limited access kind of concerns where you're trying to convey that. But there's a lot of mundane things. There's a lot of home-based things and small business-based things, as you pointed out the real estate one earlier.

12:48So this is just one technology that we're talking about so far. Yeah. And I think what you're saying, it illustrates how this is impacting very large organizations all the way down to small organizations. Yeah. So proprietorships. Yeah. And it's interesting how like if we just take this use case, for example, these kind of 3D scenes, kind of large scale organizations that maybe their bread and butter was either the compute associated with like rendering videos and 3D scenes or their hardware providers that are creating specialized kind of 3D type of equipment. like their whole business model they've got to be thinking similar to other organizations that are dealing with maybe language related problems that are thinking about these things with respect to llms there's a fundamental shift in maybe how their businesses will operate but then at the same time it provides an opportunity for the kind of small to medium businesses to embrace this technology very quickly and actually make innovative products that can be widely adopted very quickly and actually be competitors within an established market.

14:09So there's an established market for 3D things that has been quite expensive over time in terms of access to that technology. So now that whole market's going to change. I think a lot of the players will be these kind small to medium-sized businesses. I agree. I think there's a moment here, kind of ironically, because people are so worried about the impact on human creativity because of all these models and stuff like that. But on a more positive note, there's this huge opportunity that you're just now alluding to for people that if you can connect the dots as things are coming out and you can stay on top of it, it's a great equalizer.

14:47And so it will clearly change many, many markets that are out there and many, many industries. And so there's huge opportunities for those who want to surge ahead at this moment and take advantage of that. And so I think that the message we tend to see in the media tends to be a little bit doomy and gloomy on that, but it kind of discounts the fact that change isn't always a bad thing. People are afraid of it, but there's huge, huge opportunities here as well if people choose to go find them.

15:32Well, Chris, there is a new llama in town. I know. Llama 2. Llama 2 basically destroyed all of my feeds and concentration this week when it was released because it is quite, to me, an encouraging thing, but also another transformative step in what we're doing. So LLAMA 2, for those that maybe lack the context here, Meta or Facebook or however you want to refer to it, Meta had released a large language model called LLAMA, which was extremely useful. It was a model where you could host it yourself as opposed to like OpenAI. You could get the weights and host it yourself. But the original LLAMA had a very restrictive licensing and access sort of pattern, even though you could kind of download the weights from maybe like a BitTorrent link or something like that.

16:37And those propagated. Technically, if you got those weights, you were still restricted by a license that prevented commercial use cases specifically. And now with Llama 2, Meta's released the kind of follow-on to Llama and we can talk through some of what the differences are and what it is and some of what went into it. But I think one of the biggest things, which is I think going to create this huge ripple effect throughout the industry is that they've released it with a commercial license. As long as on the day that Llama 2 was released, you as a commercial entity don't have greater than 700 million monthly active users.

17:28You can use it for commercial purposes. So maybe if my company maybe later on has 700 million monthly active users, which would be great, probably never. There'll be something past Llama 2 by then, though. Yes. If it does, though, I could still actually use it because it's only on the release date. So on the release date, which was this week, as long as you didn't have greater than 700 million monthly active users, you can use this in your business for commercial use cases. And I think that's going to have a huge ripple effect downstream. And we can talk about the model itself here in a second, but maybe just I'll pause there to get your reaction on that, Chris.

18:09It made me smile when I heard that because it's kind of like saying, so long as you don't compete with us at Meta, you can use this for commercial. Oh, it's totally true. Yeah. Like, who is that? Right. So that's Snapchat. Yes. TikTok. Right. Like you can think of. Yeah, you can think of who this is. And I guess one way to put this is it's not totally open source, quote unquote. We wouldn't call this maybe open source in the kind of official definition of open source. But it's certainly commercially available to a very wide set of people. Yep. You know, one of the first things I noticed when this came out on their page and they're taught, you know, there's there's and I'm diving into like the specifics of the model here is we had an episode not too long ago and you were describing about kind of the I believe it was the seven billion limit, you know, in terms of hardware usage and stuff.

19:05And having been taught that by you, I immediately locked in on the smallest being$7 billion there. And I thought, ah, this is what Daniel has taught all of us about that limitation on accessibility and who can do it. So, you know, it has the$13 billion and the$70 billion size. But I definitely picked up on the$7 billion, which I'm assuming is going back to what you were teaching us a few episodes back. Yeah. And so just to fill in a little bit on that. So the Llama 2 release includes three sizes. So again, thinking back to what are the kind of characteristics of large language models that kind of matter as you're considering using them.

19:48One is license. We've already talked about that a little bit here. We might revisit it here in a second. Another is size because that influences both the hardware that you need to run it and also its kind of ease of deployment. So LAMA 2 is released in 7 billion parameter, 13 billion parameter and 70 billion parameter sizes. And then there's also, of course, the training data and that sort of thing that's related to this and how it's fine tuned or instruction tuned. So LAMA 2 is released in these three sizes, both as a base large language model and a chat fine tuned model. So there's the 7 ,013 ,070 ,070 ,020 Llama2s.

20:37And then there's the 7 ,013 ,070 ,020 Llama2 chat models, which we can talk about that fine tuning here in a second. But yes, you're right, Chris, in that 7 billion, I could reasonably pull that into a collab notebook and maybe with a few tricks, but with certainly with the great tooling from hugging face, including ways to load it in even 4-bit or other quantizations, I can run that, you know, on a T4, for example, in Google CoLab with some of the great tooling that's out there. So not needing to have a huge cluster. The 70 billion, even with that, that's kind of another limit where using some of these tricks, I've definitely seen people running the 70 billion parameter model on an A100.

21:30Again, loading in 4-bit with some of the quantization stuff and all that. The 70 billion is certainly going to be more difficult to run. It might require multiple GPUs, but that's kind of that sizing range for people to have in mind and how accessible things are. And yeah. How might you, I'm just curious if you're looking at these, you're a business out there or data scientist and can you make up a couple of use cases that you might target with each of these where you might say oh I want to go 13 on this not 7 not 70 for something like this can you imagine something like this I'm putting you on the spot yeah I think I mean there's certainly innumerable use cases but I think maybe two distinctions that people could have in their mind is if you want like your own private chat GPT right or like a Another way you could think about it is a very general purpose model.

22:23You could do anything with this model, any specific prompt, whatever. You're probably going to look towards that higher end, the 70 billion parameter model for that kind of almost chat GPT-like performance. You're going to have to go much higher. But as we've talked about on the show before, most businesses don't need a general purpose model. They need a model to do a thing. And so or a task or a set of tasks. And so in that case, I think businesses, because this is open and commercially licensed businesses that could take those seven and 13 billion parameter models and fine tune them for a task in their business, which also is increasingly has amazing tooling around it.

23:11again from from hugging face and others with the peft library parameter efficient fine tuning and the laura technique which is the low rank adapter technique which basically only adapts an existing model it's kind of an adapter technique rather than retraining a bunch of the the original model this opens up fine tuning possibilities in these smaller models where that fine tune for an organization is going to perform probably better than any general purpose model out there. And because it's that smaller size, you can run it on a reasonable set of hardware that's not going to require you to buy your own GPU cluster to host the thing.

23:54So that's kind of maybe a range of use cases that people could have in mind. I have one more question for you before we abandon this. 7 billion to 70 billion being an order of magnitude jump on that. Why would you have something fairly close to that at 13 billion parameters? Like what's the difference in 7 and 13 when the next step is all the way up to 70? What's the rationale you think? Yeah, so it is interesting actually if I'm understanding right from some of the sources that I've been reading. There was actually, I forget if it was 30 or 34 billion parameter. parameter model that they were also had in pre-release and were tuning.

24:37So there was another one that kind of fit in that slot that is kind of missing that gap like you're talking about. Like if you think of MPT, MPT has a 30 billion parameter model that fits in that kind of gap. My understanding and you know if our listeners can correct me if I'm wrong please do but my understanding is that they actually did test that size of model and found it to not pass their kind of safety parameters around harmful potentially harmful output or not truthful output that sort of thing so they decided actually to hold that back so it could be possible as they instruction tune and get human feedback potentially more iterations of reinforcement learning from human feedback.

25:24There may be a model that they release in that parameter range. So that was one thing that that happened, I think. It is interesting, you know, several different things here that are unique about this model specifically, or maybe the release as well, other than the license, is they were fairly vague on the data that went into the pre-training. So they talked specifically about some very intense data cleaning and filtering that they did on public data sets. And it was trained on more data than the original LAMA, but they're fairly vague on the mix of that data and all of that. So that may be related to feedback they got on the data sets that were used in the first LAMA.

26:14I don't know, but the technical paper was mostly related to the modeling and fine-tuning trickery and methodologies that they used, which was interesting. And one of those interesting elements of the way that they fine-tuned this model was, I think, the reward modeling. So if you remember, like the GPT family of models, the MPT, Falcon, these different models, one of the things that is often done with these models is this process of reinforcement learning through human feedback, which is this process, and we covered this on a previous episode, which we can link in the show notes, but actually using human preferences to score the output of a model and then actually use reinforcement learning to correct the model to better align with human preferences or human feedback.

27:09They actually use two separate reward models in this fine-tuning of the chat-based model, one that was related to helpfulness, and then the other one which was related to safety. And one of the interesting things that they talked about in the paper was how sometimes those things can kind of work against each other if you're trying to do both of them at the same time. So they actually separated out the reward models that they used for the chat fine-tuning into these two reward models, one for helpfulness and one for safety, which is quite interesting, I think.

28:01So Chris, maybe just a couple other things related to Llama. And then I want to see your feedback on Code Interpreter as well, because we haven't talked about that yet on the show. and maybe Claude 2 if we can get to it. Yeah, we got to mention Claude 2 as well because they were both big releases. Yeah, so just one maybe other note which I find quite interesting and actually I love our previous guest, Damian's thoughts on this who was in our last episode about the legal implications of generative AI. But one of the interesting things about the Lama license in addition to it allowing this commercial usage is that there is technically a restriction in the LAMA license that says you will not use LAMA materials, which includes the model weights and et cetera, or any output or results of the LAMA materials to improve any other large language model, excluding LAMA 2 or derivative works thereof.

29:02So essentially what this means is if you're using LAMA 2 and you want to fine tune a model or you're fine tuning a model off of llama to outputs, you're stuck with llama to basically llama to is your model and that you're going to stick with llama to. So you couldn't, for example, technically take the llama outputs from llama to and fine tune, say, Dolly three billion. Right. That would not be allowed by the license. And of course, that's something that people are doing all over the place. They're taking outputs from GPT-4 and fine-tuning a different model or taking outputs from a large model like, you know, maybe Lama 270 billion now and fine-tuning another model that's smaller based on a certain type of prompt or something.

29:56So this is restricting that family of models that you're allowed to do that sort of thing with, which is the first time I've seen that. I think it's kind of interesting. Yes, it strikes me as another Mark Zuckerberg anti-competitiveness thing, which he's fairly famous for. I mean, that's kind of even before this. Yeah. And how could you enforce such a thing? Yeah. That was my next question to you is, is there any possible way that you could conceive of to actually know that from an enforceability standpoint? I have no idea. I don't either. So it seems it's like it's a license thing and it will concern the lawyers, but it's hard to imagine.

30:36I mean, going back to our conversation last week, once you have output and that output is input to more output and, you know, there's a point where it becomes very, very, very difficult to know what the sourcing really was. And the fine tunes are already appearing off of Llama 2. So the most notable probably is Free Willy, which is from Stability AI and is a fine tune of the largest 70 billion model. But there's other ones coming out as well. And so I think we're about to see just a huge explosion of these Llama 2 based models for a whole variety of purposes. and who knows how they will fit into that licensing restriction or how open people will be about that.

31:24But it's about to start. The fine tunes are already coming. Yeah, well, you know, to your point earlier, they weren't terribly clear about the data that they were sourcing from their own standpoint. Yeah. And I find it interesting, a little ironic. It's a bit of a double standard maybe. Yeah, a little bit of a double standard right there in terms of like, we're not going to tell you everything about how we're doing input, but by the way, you better not use our output for your, you know, for something. So yeah, a little interesting. Do you think there's any risk of a walled garden kind of concept happening in large language models if others were to follow this lead on anti-competitiveness?

32:01Yeah, it will be interesting. I think it is a notable trend that the first llama from meta was not open for commercial at all. And now they're opening it up for commercial purposes. And, And, you know, maybe there's a separate trend that will happen with some of these use based restrictions that people are importing into their licenses and how useful those things are over time that will may shift and we'll see those things die off. Or maybe if they're enforced and there's precedent, maybe we'll see something go the other way. I'm not sure. But speaking of models that you might get their output and use it to train other models, that is these large scale proprietary closed models from people like OpenAI and Anthropic and others.

Read the full transcript

32:48We've got a couple of things that we haven't talked about on the show yet, which people should probably have on their radar. One of those is Claude 2. What do you think about Claude 2? from Anthropic. Yeah, I've been playing around with it a lot in the last week. And I kind of have a set of things that I try over and over again. They're kind of my standard tasks as new models come out. And some of them are coding and some of them are content generation, which are kind of the two big things that I use most often. It was interesting. You can put, you know, the input size for Cloud2 is much larger than the others.

33:25Like much, much larger. Much, much, much larger. So 100 ,000 tokens. Yeah. And so it's had me kind of change the way I'm approaching it in that by contrast with like chat GPT, and you're trying to figure out with, with the limits that you have both on input and output, how do you kind of prompt engineer your way to get, you know, where you're trying to go, which has become this whole skill set we've been talking about, you know, in recent months. And yet cloud two almost kind of wipes that out a little bit in some ways, not, not in all ways, and that you can hit it with a much larger input space.

33:59And, and so it's changing how I'm thinking about kind of getting to the output that I want. And the output is a bit different. It's not the same. I'm getting out different outputs from, from all the models. So yeah, they're not all the same. Definitely. I think my biggest thing is with all these new releases, I'm trying to figure out how do I use each one? When do I, I'm trying to develop my own strategy on when do I go to chat GPT by default? Like when's that the right thing? And that's changing as we'll talk about with things like plugins and stuff that's evolving. But then Claude 2 comes out and then you have, you know, on the open source side, as we just talked about with Llama 2.

34:33So I think trying to understand all the tools in the toolbox in relation to each other has been interesting. So Claude 2, I'm really focused right now primarily on large content output is kind of where I've landed on that. And the 100K context length of Claude 2 is something I find really compelling as well. There was also a significant paper that came out that caused a lot of waves in terms of context length and thinking about that, which showed kind of as you increase context length, you lose any significance of the middle bit of that context. So the beginning and end is more important in terms of what makes the output of the model quality or not in terms of how you would measure that.

35:22And so we'll link to that paper maybe in the show notes as well. But I've tried some things. I mean, I don't know exactly all of the details. Again, Claude is one of these closed models. So I don't know all of the details of how they're doing things. And because it's sitting behind an API, it's hard to know how those things evolve over time. But for example, I took one of the things with Cloud2 is I just took one of our complete podcast transcripts. So a full episode, so 45 minutes of audio transcript. I took episode 225, which I really enjoyed talking a lot about the things that I'm working on right now with Prediction Guard.

36:01and just asked it to give me a summary of the main takeaways and, you know, paste it in the whole thing. And it's like a fairly good comprehensive takeaways. Like many companies ban usage of certain LLMs, blah, blah, blah. You know, Prediction Guard is trying to provide easy access, structuring validation, compliance features for LLMs, making LLM usage easier, blah, blah, blah. And it gives these great takeaways. And then I asked, you know, hey, suggest a few future episodes that we could do that maybe cover related topics, but things that weren't covered in this episode. Pretty good. Some of them are kind of generic, right?

36:44I'll look at current state of AI agents and automation. How close are we to no code AI app generation, blah, blah, blah, blah. So that all kind of off of this large context of the transcript input was quite interesting. I'm curious. I'm going to put you on the spot also. As someone who's working on your own product, and I know this is not a Prediction Guard episode, but I'm asking on my own behalf and on behalf of the listener, how do you, as someone who is looking at these different models, how do you think of those different models? How do you kind of structure them in your mind in terms of what you're offering?

37:18You've been evolving rapidly over the last few months, and I'm always curious to see kind of where your head's at on this now as you're looking at them. Yeah, I think the things consistently that I'm seeing are that I made a post on LinkedIn about this as well. Even my own applications that I'm building, LM based applications, having access to multiple models rather than a single model, I think is a really nice usage pattern where if The easier we can make it, and there's other people that are doing this as well. In PredictionGuard, you can query a whole bunch of models at the same time concurrently.

37:57There's other systems that will let you look at that output as well. So NAT.dev and some of the toolbar stuff that SWIX is doing. We had a collaboration with him in the Latent Space podcast. So the more you can tie these things together and look at the output or automatically analyze the output of multiple models at the same time, I think that's really useful because it's hard to generally evaluate these models until you start evaluating them for your use case and building intuition about them for your own use case. So I think the pitfall that people maybe fall into is saying, oh, I'm going to use this model before they've even tested that for their use case.

38:39Try creating a set of evaluation examples for your own use case and then try out a bunch of different models for that. And also try out the things that are becoming more standard kind of operating procedures for building LLM applications, like looking at the consistency of outputs, running a post-generation validity or factuality check on the output. But so checking a language model with a language model, doing input filtering and all these sorts of more engineering related things. So those are some of the things that I'm seeing. But having access to a bunch of models at the same time, I think, is something that can really boost your productivity.

39:25I appreciate that. And to our listeners, we're not making it a prediction guard show or episode. But as a co-host, Daniel's excursion through this and his professional has made him, in my view, one of the world's true experts in how to look at all these together. And since we have the benefit of him co-hosting the podcast, I'm going to continue to take advantage of that expertise for all of us. Thanks, Chris. Sorry about that, Daniel. Sorry for putting you on the spot. Yeah, no, no worries. I think the other thing maybe to highlight with Cloud2 and something that you were talking about in chat before we jumped into this episode was Cloud2 versus or maybe Anthropic and their offerings versus OpenAI.

40:08How do we understand that? Like how do we categorize these things? I think one of the interesting things with Cloud2, so we've seen both Anthropic and their Cloud models and OpenAI and their GPT models increase context size over time. GPT models, not quite as far as Cloud, but both have increased. They've also both added in some of this functionality, which I think is very interesting. Cloud2, I think, first, if I'm not wrong, the ability to add in your own data. So in Cloud 2, there's a little attachment button and you can upload PDFs or text files or CSVs and have that inserted into the context of your prompt, which I think is, of course, extremely powerful.

40:54We've talked about adding in external data into generative models and grounding models in the past. It's very powerful. Now, OpenAI is doing this in a slightly different way. and I think this is something worth calling out on the podcast is with their new code interpreter beta feature within ChatGPT, you can upload data, but it's processed through the code interpreter in a different way than what Cloud is doing. So we all know that ChatGPT and GPT models can generate really good code and specifically good Python code. And so what OpenAI has done for their kind of data processing agent within chat GPT is said, well, let's just have our model generate Python code.

41:44Then we'll hook up the chat GPT interface to a Python interpreter and just go ahead and execute that code for you over your data and then give you the output. So this is maybe a distinction that people can have in their mind. Claude 2, you can upload this huge amount of context you can upload files insert it into the prompt as far as i know they're not running any kind of code interpreter type thing under the hood chad gpt might not be inserting all of that into the prompt but they're actually saying well what if we decompose what you're wanting me to do with this external data into something that can be executed by a sort of agent type of workflow where you upload your data and ask me to do some analysis over it.

42:33I'm going to generate some code. So the language model generates some code and then that code is actually executed in the background. It returns a result, which is then fed back through a model to give you generated output back in the interface. So it's actually a multi-stage thing happening in Code Interpreter in OpenAI. It effectively produces a no-code solution where you get an output and you're just kind of skipping the whole thing. Instead of using the language model to generate your own code and to be your code assist and all that, and then you're still doing it, it's kind of skipping that whole step right there.

43:09Yeah, and I can give an example I actually ran prior to the show. So I have Claude and the OpenAI code interpreter side-by-side open. I uploaded a file with a bunch of Yoruba, which is a language in Africa, transcriptions out of audio, which are from the Bible TTS project that we worked with Koki and Masakane on. And so I uploaded this file, which includes this Yoruba text in a CSV format. OpenAI said, great, you've uploaded this file. Let's start by loading and examining the context. And then it has this sort of show work button. And you can see the actual code that it generated, which is Panda's code to import the CSV and then output some examples.

43:58And so you can expand that and actually see the code that it ran under the hood and the conclusions that the agent came to. Then I asked it, OK, well, plot the distribution of the transcript links. Are there any anomalies? And then again, it says, hey, show work. And you can see it's importing matplotlib. plot lib. It's taking in the CSV. It's actually creating the plot and actually generates an image out of the transcripts. It says, I didn't find any anomalies. They're all kind of within the same distribution. There's not any anomalies. Then I asked it, can you translate all the Yoruba to English?

44:33And that's where it ended up stopping because it said, no, I'm not good at doing that. And Quad actually stopped there as well and said, no, I'm not going to do that. I also uploaded the Yoruba alignments to Claude and it said, hey, sure, let me analyze these transcripts. And it just output some general like there are 50 audio links, the transcript links. There's no Python code there. It just gave me some takeaways. Right. And then I said, are there any anomalies? And it said, I checked and I can't find any. And could you translate it? And it said, unfortunately, I can't. So it's all still a chat based thing.

45:10So you can see kind of different approaches to this complicated workflow of having almost an assistant agent executing code for you versus putting more context in the language model and having it reason over that context. So they're almost getting their own strengths at different types of approaches to problems. Would that be fair? Yeah. So that's another way of thinking about it is as you start understanding how the different large language models approach a problem and the tooling that might be better or worse for a given use case, that also will help you kind of pick which way you want to go.

45:46In addition to maybe just using multiple models, as you talked about earlier. Yeah, exactly. And there's so much to dive into on all these topics that we've covered today. I'm going to make sure that we include some really good learning resources for people in the show notes. So make sure and click on some of those. There's a guide from Datagen on the neural radiance field stuff, the Nerf stuff that you can learn a bit more about that. there's a hugging face post and a Phil Schmidt post on Llama 2 that are both really practical kind of how do you run it how do you fine tune it what does it mean and then there's a nice post from the one useful thing Ethan Mollick blog or newsletter about code interpreter and how to get it set up and some things to try so we'll link that in our show notes and I think people should dig in, get hands on with this stuff.

46:40Things are updating quickly. And the only way to really get that intuition about things is to dive in and get hands on. It is. It's the most interesting moment we've had in the AI revolution of recent years and just so much cool stuff right now. Anyway, thank you for taking us through all the understanding and explanation of these things. Yeah, definitely. It's a good time. Hopefully people enjoy the rest of their week and maybe go see Oppenheimer or Barbie, depending on your, which of those is most interesting to you. But we'll see you next time, Chris. See you later. Thanks.

47:35and Fly for partnering with us to bring you all Change Dog podcasts. Check out what they're up to at Fastly.com and Fly.io. And to our Beat Freakin' Residence, Breakmaster Cylinder, for continuously cranking out the best beats in the biz. That's all for now. We'll talk to you again next time.

48:08Game on!

From the publisher

It was an amazing week in AI news. Among other things, there is a new NeRF and a new Llama in town!!! Zip-NeRF can create some amazing 3D scenes based on 2D images, and Llama 2 from Meta promises to change the LLM landscape. Chris and Daniel dive into these and they compare some of the recently released OpenAI functionality to Anthropic’s Claude 2.

Join the discussion

Changelog++ members save 1 minute on this episode because they made the ads disappear. Join today!

Sponsors:

  • Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com
  • Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs. 
  • Typesense – Lightning fast, globally distributed Search-as-a-Service that runs in memory. You literally can’t get any faster! 

Featuring:

Show Notes:

Learning resources:

Something missing or broken? PRs welcome!

More from Practical AI

All 157 episodes
There's a new Llama in townPractical AI · 48 min
Listen in VO