Udio & the age of multi-modal AI

16 Apr 2024 · 39 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Practical AI Podcast Episode Notes

Episode Title

Udio & the Age of Multi-Modal AI

Episode Overview The episode dives into the emerging trend of multi-modal AI, with a focus on a new music generation service called Udio. Hosts Chris Benson and Daniel Whitenack explore the transitions from traditional data modalities in AI to the innovative integrations seen in multi-modal systems.

Key Themes

  • Emergence of Multi-Modal AI: The discussion emphasizes that 2024 is expected to be a pivotal year for multi-modal AI technologies which integrate various forms of data inputs and outputs.
  • Udio Music Generation: A significant portion of the episode is dedicated to Udio, highlighting its capabilities in generating music, lyrics, and even singing voices from text prompts.

Hosts

  • Chris Benson: Principal AI Research Engineer at Lockheed Martin
  • Daniel Whitenack: Founder and CEO at Prediction Guard

Episode Highlights

Opening Remarks

  • Daniel and Chris express excitement about the rapid advancements in AI, including new models like GPT-4 Turbo and Gemini.

Multi-Modal AI Predictions

  • 2024 is predicted to see an explosion of multi-modal AI, moving from primarily text-to-text applications to more complex integrations of text, image, and sound.
  • They share a lighthearted moment about how their discussions might be influencing trends in the industry.

Udio Overview

  • Udio is introduced as a platform for generating music based on user prompts.
  • The hosts provide a demonstration of Udio’s capabilities, including an AI-generated song titled "Dune: The Broadway Musical".
  • The platform allows users to input various stylistic elements to influence the genre and mood of the generated music.

Music Generation Demonstration

  • The hosts play several AI-generated music samples, showcasing the potential for creative outputs and the ease of generating complex compositions quickly.
  • Discussion on the implications of AI-generated content in relation to copyright and creativity, suggesting current laws may not adequately cover AI-generated works.

Legal and Creative Considerations

  • Exploration of the legal grey area surrounding AI-generated content and the potential for changes in copyright laws as multi-modal AI becomes more prevalent.

Historical Context of Multi-Modal AI

  • The hosts provide a historical overview of AI development, from specialized models for text, speech, and vision to the current integration of multi-modal capabilities.

Future Directions

  • Discussion on how advancements in emotional recognition and personalized content generation could expand the applicability of multi-modal AI in various domains.

Key Takeaways

  • Udio’s Impact: Udio demonstrates a significant leap in AI music generation, combining lyrics, composition, and vocalization in a single platform.
  • Legal Evolution: As AI technologies evolve, existing laws regarding copyright and creativity may need to adapt to new realities.
  • Encouragement to Experiment: The hosts encourage listeners to engage with multi-modal AI technologies through hands-on experimentation, such as using Udio or exploring models like GPT Vision and LLaVA.

Recommended Resources

  • [Udio](https://www.udio.com/)
  • [CLIP by OpenAI](https://openai.com/research/clip)
  • [BridgeTower](https://arxiv.org/abs/2206.08657)
  • [LLaVA](https://llava-vl.github.io/)

Conclusion Chris and Daniel wrap up the episode by reinforcing the importance of staying engaged with multi-modal AI developments and the creative possibilities they present. They express excitement for the future of AI and encourage listeners to explore these advancements further.

---

Subscribe and Join the Community

  • Subscribe: Join the Practical AI community at practicalai.fm.
  • Join the Slack Community: Engage with the hosts and fellow listeners [here](https://practicalai.fm/community).

Sponsors

  • [Fly.io](https://fly.io/changelog): Services for deploying applications close to users.
  • [Changelog News](https://changelog.com/news): A podcast and newsletter for developer news.

---

This document serves as a comprehensive overview of the discussed episode, outlining key themes, takeaways, and resources for further exploration in the realm of multi-modal AI.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related tech is changing the world, this is the show for you. Thank you to our partners at Fly.io, the home of changelog.com. Fly transforms containers into micro VMs that run on their hardware in 30 plus regions on six continents. So you can launch your app near your users. Learn more at Fly.io.

0:42Welcome to another fully connected episode of the Practical AI podcast. In these fully connected episodes, Chris and I keep you fully connected with everything that's happening in the AI world, the news, the trends, the new models, all the good stuff, and talk through some things that will hopefully level up your machine learning game. I'm Daniel Wynack. I am founder and CEO at Prediction Guard. And I'm joined, as always, by my co-host, Chris Benson, who is a principal AI research engineer at Lockheed Martin. How are you doing, Chris? Doing great today, Daniel. I don't know how we're going to pick what to talk about.

1:21There's so much stuff coming out right now. There's a lot. Mostly kind of new. Well, there's always new models, I guess. but it did seem like a big week with, I think, new GPT for Turbo, new Gemini. It's really hard for me to keep track of the numbers and parameter counts and all that. But I know that new Gemini, I had 1.5 something. I forget the different numbers. And then new Mistral, new, I think it's a different mixture of experts, top of the open LM leaderboard. We've got Udio, which I was just so I've been at some startup accelerator stuff and then at conferences, events. And I'm like, it seems like I go into one of these events.

2:09And at the end of the day, people just say, oh, did you hear about X? And I'm like, no, I was I haven't had my laptop open. Something else has happened, apparently. For the last two hours, I haven't had my laptop and I've missed I was doing work, you know? Yeah. I don't know if you've seen the trend. I think this was a prediction for 2024, which I think was a well-informed prediction for 2024 from many different people. And I think we talked about this in our own discussions of 2024 around multi-modality of AI in 2024. Whereas in 2023, you kind of saw this explosion of, in many cases, text-to-text AI, meaning I put in a text prompt and I get some text back.

2:57Now we're seeing an explosion of multiple modalities of data input and or output to these models. and that's mostly what I'm seeing. Is that consistent with your view as well, Chris? Totally. I was thinking about that and a moment for us to brag. I tell you, we've actually been fairly good with our predictions the last few years. Who knows? Maybe we're actually setting the, it's a self-fulfilling prophecy. It's just everyone's listening to practical AI and they're making our predictions real. That must be what it is, I'm sure. You know that all of the, you know, like over at OpenAI, they're just listening to us and they're going, okay, that's what we need to go work on.

3:37It's a lot of pressure. Yeah, I tell you what, it's a good thing we're steering the entire industry by ourselves, right? Yeah, it's a lot of pressure, but it's good. Sure. No pressure at all, man. I do want to talk about multimodality today, but I've just got to share with you some of this Udio stuff, Chris, which I think, well, Udio, Udio, I don't know if anyone knows how to say this yet. I was saying Udio, maybe it's U-D-O, U-D-I-O, I believe. I'm not going to hazard a guess until I hear somebody else do it. Yeah, yeah. It's coming out of beta. I think that there were some leaks of some of what they were doing before.

4:18But essentially, if you go to this website, you can sign up for an account. They have it marked as beta. So I'm not sure exactly where this is going necessarily product wise. But what you see, at least in its beta form, is essentially a space where you can put in a text prompt. It kind of reminded me of almost like a clip drop or something. Some of these image generation platforms where you can kind of pre-select some elements of the prompt. And the goal would be to completely generate a coherent and compelling song or piece of music or composition. it's essentially a music generator so we've seen a little bit of this in the past right and in the past we've kind of heard things generated like kind of dreamy ambient things and and maybe useful for kind of backing youtube videos or something like that but not really compelling music in and of itself.

5:25And I think what's interesting about this UDO is that it generates both this sort of compelling music, but also lyrics and also synthesized voices singing the lyrics all together in one. So Chris, before, while you were setting up your studio there to record the podcast, I was busy on UDO figuring out what is there. Now there's a couple of really interesting ones that I listen to. And I've preloaded a couple in here for us to listen to and to give our audience a little bit of a sense of the audio. Absolutely. Because this is an audio podcast. So, you know, what better format? So one of these, which I found really intriguing, was Dune, the Broadway musical.

6:16And I would go to that, by the way, just to make it very clear, I'm standing in line to buy tickets, so to speak? Well, the music has been generated for this and I'll cite Bobby B. So Bobby B on UDO and he's created Dune the Broadway musical. So just to give a sense of people, the prompt that went in to this to create it, it says teen pop show tunes, film soundtrack, uplifting, playful, female vocalist, happy. Anyway, so you get a sense of kind of similar to a, I guess, like an image generation prompt where you're saying like, high resolution, Unreal Engine, like this sort of stuff to give it some stylistic guidance.

7:02But I've got this preloaded in here, Chris, everything that you're going to hear is AI generated. So let's listen to Song of Arrakis from Dune, the Broadway musical. Oh, you got my attention, man. Paul, trade deeds of Arakeen The greatest leader we've ever seen They say that he's the list in Al-Gaib Eyes bright blue and hair jet black You should see him ride on a sandworm's back Lead us to victory, you soul Mu 'ati What do you think? Move over, Wicked move over les miserables all just you know i saw hamilton recently move over hamilton yes yes we're all about doing the musical now that's it exactly yeah so good and even the lyrics you know eyes bright blue and hair jet black you should see him right on a sandworm's back i mean that's great that's good right there i like the fact that the music actually deviated far from the darkness of Dune, you know, the perpetual darkness of the theme.

8:09That was fun. That was great. All right. Ready for more? Yeah. Oh, yeah. So I tried out my own. Of course, I had to try out my own. Of course. So my prompt that I put in, and I only experimented with this, so I'm sure you can do much better. But I put in a song about two podcast hosts trying to navigate the wild and crazy world of AI in the style of pop rock. Practical AI, the musical. We really appreciate you, Breakmaster Cylinder, the mysterious Breakmaster Cylinder behind our theme music for the show. But this is what UDO can do. And I selected specifically to have it pop rock and to auto-generate the lyrics.

8:56So I didn't put in any lyrics. So I have two options for you, Chris. I have two selections that you can see which one you like better. Here's the first.

9:30that's as quick as a flash. It's a step forward. It's a step forward. All right. Selection one. Thoughts? I like that one. You just transported me from like, you know, 53, which is what I'm at now, all the way back to like 16 in the 80s, you know, late 80s. I was all about that. That was good. Okay, cool. Yeah, I love it. Okay, that one was, they also generated a title. It's interesting. I don't know how much of this, you know, how many models are at play under the hood here and how they're coordinated. I'm guessing maybe there's some that generate the lyrics and some that generate the title. And then somehow that's merged together in a music generation.

10:13Because obviously the voice and the lyrics have to be coordinated somehow. I at least didn't see a lot of underlying explanation of what's going on here, but pretty interesting. And that was generated, I would say, in 30 seconds or something. I don't know. Not that long, right? So let's take a listen to number two. This one was titled Digital Odyssey.

11:02There you go. Even with a little bit of a guitar type solo there at the end. I know it was good. I like the first one better. The first one felt like I was a kid again right there, but I like both. And yeah, this was good. I could just spend all day generating music now. I may do that actually. I might have just taken up your Saturday. Oh my God. My wife has all these chores planned for me because we're recording here late on a Saturday morning and I may get myself in trouble by. Yes. Okay. Well, there you go. So UDO, check it out. Super cool stuff. I think this does bring up some really interesting challenges, issues, struggles, excitement.

11:49Also, of course, joy in hearing Dune the musical. But it is super interesting. I even thought about this when I was going through here. Well, you know, Bobby B created Dune the Broadway musical, and I just downloaded and I'm doing it what I want with it, which I guess is playing it on my podcast. So Bobby B, I hope you're okay with that. Now, technically, technically, this is machine generated. So at least as far as I understand, in the current US legal system, such a thing would not be copyrightable. Sorry, Bobby. Sorry, Bobby. I'm not giving legal advice here, obviously, and not a lawyer. But that's my understanding from our previous conversations.

12:38But what's interesting is I think similar to these AI generated art things that, you know, were put into art competitions and won, right? There could be now, in my case, I just put in a simple prompt that, you know, generated something in 30 seconds. There could be some really deep thought put into how to construct this prompt and the various kind of, I think Bobby's prompt was much better constructed than mine. And also you can upload your own lyrics into this to add a level of creativity. So there's really an open question here of how much human creativity is actually a big portion of this generation.

13:26And will the established laws and legal entities eventually recognize the creativity that's put into prompting these sorts of systems. Just like there was a time when there was a question of whether or not if you took a picture with a camera, right, you just click the button. Now, photographers out there are going to get really mad at me because I think they would recognize it's way more than clicking a button, right? There is a whole lot to photography. And that's, I think, why it's been accepted as an art. But people argued at one point, you know, hey, you click that button, it's machine generated, you can't have a copyright for that.

14:09But eventually that those laws change. So I wonder, Chris, if you have any thoughts about it, if or when that might change in these cases? Yeah, I mean, I think it will, because I guess if you look at these stream of AI advancements that we've covered over the years, there's a sense of inevitability that when these things come out, they catch on and they become popular and then they become the norm. And then eventually, as we keep seeing the laws gradually catch up over time and, um, things like UDO, if I'm pronouncing the name right, is going to be typical. It won't just be them. There'll be others as well.

14:46And so I know that, you know, by way of example of that inevitability, we have a Spotify account for our family and, um, you know, we're listening to the traditional way of streaming music historically. And one of the things I do on that is I really like to explore new genres and new types of music that I don't know. And I'm always trying to think, how do I get to that? But I'm very likely to use something like UDO to prompt what I'm feeling, what I'm thinking, and try to explore new music that way. Because I don't really care as a user whether or not it's an artist that's human or an AI model that generated if it sounds good to me.

15:22And so I think that sense of inevitability will bring about the change over time. Yeah, I think personally, I think right now, even it's a gray area where, especially if you're going back and forth, like you're trying a prompt, and then maybe you're modifying the lyrics. And if there's some sort of back and forth, that definitely gets into a little bit of a gray area where how much even of the generated stuff, ignoring the creativity and the prompt, how much of the generated stuff is actually machine generated versus human post edited, for example. That's right. So yeah, I think that even now that's a bit of a gray area, but then my personal thought is eventually this will be more recognized as a creative pursuit, but you know, we'll see.

16:11You know, this new inroads into music through this AI model with this being so far the most interesting that I've seen, this probably will really scare the music industry, you know, because this is taking it to a whole different level. And there probably will be a lot of lobbying, a lot of lawsuits. You know, we saw this past year, you know, actors going on strike because of AI based video and the creation of characters or the representation potentially of live people. And I think we'll see some form of that here. This is a process we're going to go through over and over again. And I was talking to a good friend just the other day about this and life ahead and stuff and how to do this.

16:54And I said, the smart people will align themselves with these capabilities. It's not about whether it's a good future or a bad future or whatever from perspective, but it's an inevitable future. If I could give advice as a non-lawyer and non - professional musician in the music industry, but someone observing this, I would say, find a way to get on board with it and make it work for you quickly because it's not going away. yeah yeah very true and i do also wonder of course those things that we just heard were completely ai generated but it's interesting to me that maybe a creative person who is and there are many that are embracing some of these things like musicians could actually iterate very very quickly on different ideas putting their own voice to backing music or getting prompted with lyrics that aren't quite so good as what they would like, but gives them a creative starting point and really explore spaces that they might not have explored before.

18:00So that might be cool to see as well, that kind of human UDO teaming. Yeah, I agree. And another thing that I think will become inevitable, so here's a startup idea for folks, is with all the advancements over the last few years in kind of emotional recognition from models and understanding if you combine a capability like this and you choose to opt in, which there's privacy concerns, obviously, with a service that also is monitoring yourself. And maybe the data is only available to you, but can generate content that is exactly specific to what you're dealing with in life. And when you need to pick me up, not only does it find the right music, but it finds the right lyrics for the situation and stuff.

18:49And so there's a lot of interesting psychological considerations here that could be both good or bad, obviously. So, but I think that's pretty fascinating. I'm wondering if I can find a service in a few years that will, that will do that. And it follows me through the day and I, I keep the content private to me in my account, but I can, it gives me the pick me up and when I want, that's what I'm looking for, for whoever is going to go out and do that in the world. Personal soundtrack and narration and vibe. My life. Yeah.

19:34This is a changelog news break. YouTuber Internet of Bugs posted a lengthy breakdown exposing Devin's creators, Cognition Labs, for falsifying claims about their world's first AI software engineer. Devin was pitched as a fully autonomous software developer, and one of the more impressive demos showed it completing and getting paid for freelance jobs on Upwork. Sound too good to be true? It did to Internet of Bugs, who says, quote, I broke down the Devin Upwork video frame by frame, and here I show what Devin was supposed to do, what it actually managed to do instead, and how bad a job of that it did.

20:17On the whole, that's not surprising given the current state of generative AI, and I wouldn't be bothering to debunk it except 1. The company lied about what Devin could do in the video description, and 2. A lot of people uncritically parroted the lie all over the internet, And three, that caused a lot of non-technical people to believe that AI might replace programmers soon. End quote. Devin really did garner a lot of attention, also known as money, because of that demo. We talked about it on our shows with a healthy amount of skepticism, I think. But I'm thankful their claims have been debunked.

20:50And I hope we all give Cognition Labs the side eye from here on out. Exaggerating your development capabilities? Maybe Devin really is human after all. You just heard one of our five top stories from Monday's Changelog News. Subscribe to the podcast to get all of the week's top stories and pop your email address in at changelog.com slash news to also receive our free companion email with even more developer news worth your attention. Once again, that's changelog.com slash news.

21:27a lot of these things like i say are moving into this multimodal sphere and it might be worth just kind of looking back a little bit at how we got to where we're at in terms of multimodal functionality, sort of how that gradually has changed over time from NLP and speech to multimodal models that we're seeing now. I think one good way to, if we kind of step back and look at it on a holistic or historical standpoint, you kind of started out with modes of data processing that were maybe separated, but often tied together in a sort of chained way. We didn't really think about it chaining at the time, right?

22:18But you had speech synthesis models, for example, that were really specifically trained to only do text to speech, right? And in some cases, even that was broken up into sub models of like a vocoder and other types of models. And you had text to text models, you had maybe computer vision models that would process images to do object recognition or even videos in certain cases or frames of videos. But all of these were specializations. So the whole idea of there being computer vision as a specialization is that I am specializing in models that process this mode of data. And speech technology, the discipline of speech technology is a discipline of really focusing on processing either speech inputs or speech outputs.

23:19And then NLP, quote unquote, had special models that would take in text and maybe classify or detect entities or do machine translation or these sorts of things. So we kind of that historically was kind of how the field was developing. And if we skip kind of the middle portion and come back to it, now we've gotten to a point where there's seemingly these large foundation models that are able to take in multiple inputs at the same time of multiple modes. So for example, an image and text in the same input paired together and, you know, answer questions. So this would be like what we see with GPT vision, or we see a text prompt or even a video input like we have seen with Gemini recently, where you can import a whole video and ask for a summary of all the visual components and that sort of thing.

24:21That's kind of, from my end, how I view the bookends. Is that also, from your standpoint, any comments on that, Chris, in terms of how we've progressed from one end to the other of that? I'll pivot slightly in response to that and say, as you were describing that, it really resonated with me on, uh, with something else that I've been thinking in that development. And that is, um, I consume a lot of content, uh, through audio books. Anytime I'm kind of on autopilot, you know, driving or mowing the lawn or doing any kind of thing, I'm listening to audio books, uh, for learning purposes, mostly at kind of a double speed as fast as my brain can process it.

25:01Cause I like to consume as much as I can. I can't do that. My mind is too slow. No, you get used to it after a while, but it's just to get the information in. And I just went through a really Pulitzer Prize winning book through audio called An Immense World by Ed Young, which is fascinating. And I highly recommend it to anybody that wants to do it. But it is all about the way we and all animals and not only humans, but all animals in their unique ways, perceive the world through their senses and how vastly different those are. And the theme that came up to me throughout that was how multimodal everything about humans are.

25:39The way that we learn, our experiences are all multimodal. We don't have just vision and just audio and just text. We're taking it all in at the same time. I think this progression that we've seen in terms of moving into multimodal this year has been really fascinating in terms of coming in to really how we take in information and how we learn. And I think going back to UDO today and seeing what they're doing and looking at the other multimodal capabilities that we've been learning it feels like we're finally getting to some i know we keep saying this it's always kind of cool at the moment when the new thing comes out but it feels like it's really aligning with what it means to be human as well ironically that's the background thought process i had as you were going through that yeah there's these scenarios in which knowledge is definite like we process knowledge across multiple modes of data inputs.

26:31And certain things are not all, many things are not all represented in text or in any given mode. And I think you've seen this already kind of utility over this with things like GPT vision, which is a kind of visual instruction tuned model. And maybe that's something to share with the audience. If you're not familiar, there's kind of this the music generation stuff maybe that's a little bit newer but there's kind of this ongoing work in visual instruction tuning and this would be the type of model in which you would have an image input and maybe a text prompt and traditionally Like I remember, I think there's even some of these models still that are quite popular to use on, for example, AWS Textract, for example, is a OCR system, but you can also do visual question answering.

27:35Now, it used to be you had a specific model architecture for visual question answering. It was a research topic in and of itself. There was a specialized model. And this kind of illustrates some of the progression that we've had. There was a very specific discipline around visual question answering and very specific models that could do those things and they advanced. But then recently you've got what has begun being termed visual instruction tuning for models where the models are actually similar foundation models to what people are using for other modes. So for example, if we look at the LAVA model, L-L-A-V-A, so not llama, but lava, maybe a bit hard to distinguish in the audio.

28:25That's a open source manifestation of the GPT vision system or similar functionality to that. And if we look at how that operates, it actually is built off of, and this is kind of, we talk about this a lot on the podcast, Chris, where you're always sort of building on the shoulders of giants and a lot of what's come before, even though some of these functionalities seem to pop up out of nowhere. But there are kind of previous signals. And Chris, I don't know if you remember, we, I think, had an episode where we talked about Clip. Yes. Which was a multimodal way to embed both text and images in something developed by OpenAI.

29:10Contrastive language image pre-training is what Clip is from OpenAI. Correct. Which, thankfully, is open to everyone back in the days when OpenAI was open and we can still use it. But Clip allows you to embed an image or text in a similar embedding space, which means you're converting an image or a piece of text into a set of numbers. and if you compare those sets of numbers in that vector space you can actually find things that are semantically similar by the distance between those vectors which is interesting and kind of makes immediate sense if you're doing text-to-text things like the semantics of one piece of text or the meaning of one piece of text could be similar to the meaning of another piece of text.

30:01It's very intriguing though if you make this multimodal and say a nice sunset on a beach in Florida and then you have an image of a sunset somewhere on a beach and then you have an image of a car driving through New York and then you have an image of a spaceship in outer space and you could actually find which of those images is semantically similar to a text input right so that's kind of the clip way of embedding things. And then on the other side, you have large language models, right, which can take a text prompt, reason over that text prompt, even though it's not really reasoning, it's just auto completing, but we can think of, you know, functionally, it takes in that prompt and outputs some output related to the question that's input, the query, the instruction, that sort of thing.

30:54So what they've done with Lava, which has been around for some time and people have built different types of Lava models and sort of its own family in and of itself, is paired the clip style embedding model or a visual encoding system with a large language model and then created this text and image input. So if you look at the architecture of what they do, what happens is they have an image input that goes through the vision encoder, for example, clip that produces an embedding. They have a language model like Lama that accepts a language instruction or text input, and that creates an internal hidden representation embedding.

Read the full transcript

31:45and the first thing they do is they train a projection matrix for the vision encoder. And can you talk about what a real quick what a projection matrix is? Yeah, yeah. So the language model produces an embedding embedded representation of the text. The vision model creates an embedded representation of the image. But these two are different model architectures and the embeddings can't be directly compared one to another because one's llama and one's clip, even though they functionally produce embeddings. So the projection matrix is a sort of translation of the output of the vision encoder, the clip model into a space in which it's concatenated or combined together in some way with the output of the llama model.

32:34And that is a trained projection such that it accomplishes the end tasks that you're training for, like a visual question answering or reasoning over an image, that sort of thing. So that's the initial pre-training is finding that projection matrix. And the interesting thing here is actually this, it's a combination of models, which is intriguing, right? Because you can always update Llama to the next cool thing like Gemma, or you could always update Quip to the next cool thing like Bridge Tower from Intel and combine them in really interesting ways and do this retraining. And then people then fine tune these models based on data sets that they've created for specific tasks, like we've seen with language models.

33:19So there might be a science question and answer for reasoning over science images, right? Or sort of visual chat in a specific domain. So to give people a sense of the functionality of this type of model, if you haven't played around with GPT Vision or something like that, one of the examples on the lava paper site, which I find interesting is there's like a meme image of a world map, but it's made out of chicken nuggets. And so it looks like a world, but it's made out of chicken nuggets. So the picture is there. And then the user input, the text input, along with the image input is, can you explain this meme in detail?

34:02So there's some element of the question that's needed to answer that question because you're saying this is a meme. You're asking for specific details, right? And then you definitely need the visual content to answer that question. Otherwise, you would just hallucinate something about a meme. Sure. So the lava answer is the meme in this image is a creative and humorous take on food with a focus on chicken nuggets as the center of the universe. The meme begins blah, blah, blah. And it essentially explains the humor, which is maybe not the best way to make something more humorous. Yeah, a little bit dry there.

34:42Yeah. But you can think of other cases where you would need sort of visual input and text input to create an answer, right? Like if you say, what was the guy who raised his right hand in this video wearing? Right. That necessarily like as a human, we would process that both from the text input standpoint and the visual content standpoint. And so I think it's really interesting, this sort of exploration of not just chained models from multiple modes together, like we've seen in the past, kind of in history where you You have a speech model, you have a computer vision model, you have a language model, and you can chain them together in interesting ways.

35:24But this joint encoding, this joint processing of multiple modes of data at the same time is actually required for some of the types of reasoning that we might want to augment or automate from a standpoint of how we process information as humans. So, yeah, I would recommend that people look into this Lava model. It's open. There's multiple, like I say, it's sort of a family or a style of doing things. So there's a bunch of examples of that on Hugging Face and demos that you can actually try out and try kind of an open version of what you get with GPT Vision. That sounds good. I've just been struck through our entire conversation, going back to what I mentioned earlier about how close this is in terms of matching how we as humans process.

36:16As you took us through the merging of the modalities a few minutes ago, and me having just, I'm actually in the middle of a, I think it's called The Great Courses, and I'm about the best brain. And it's talking about how exactly that happens in our brain to convert it into a form, which is chemical, you know, electrical in nature for our brains to actually operate on since we don't actually see and smell. And so it's just fascinating that while the underlying AI models are not the way the brain operates, that kind of the modalities are starting to merge in that way. It's really neat. So thank you very much for taking us through that understanding of how multimodality works in a practical sense.

37:01You lived up to practical AI in all ways there. Well, this episode has a little bit for everybody where you get a fun Broadway song. And then we also talk about projection matrices. So there you go. Something for everyone. I enjoyed it, Chris. And yeah, I think this will be a trend that we continue seeing throughout this year. So if you haven't got hands on and tried a little bit of this multimodal stuff, whether you go to UDO and try to create a song or you go to chat GPT and try to use GPT Vision or Gemini and process a video or download the lava model and try to run some multimodal queries.

37:39That's the best way to sort of get an intuition for how these things behave and what's possible. We'd really encourage you to get get hands on. So homework assignment between now and next episode, I guess. Absolutely. All right. It's been fun, Chris. We'll talk to you soon. Take care, Daniel.

38:04All right. That is Practical AI for this week. Subscribe now. If you haven't already, head to practicalai.fm for all the ways. And join our free Slack team where you can hang out with Daniel, Chris, and the entire ChangeLog community. Sign up today at practicalai.fm slash community. Thanks again to our partners at fly.io, to our Beat Freakin' Residence, Breakmaster Cylinder, and to you for listening. We appreciate you spending time with us. That's all for now. We'll talk to you again next time.

38:49Game on!

From the publisher

2024 promises to be the year of multi-modal AI, and we are already seeing some amazing things. In this “fully connected” episode, Chris and Daniel explore the new Udio product/service for generating music. Then they dig into the differences between recent multi-modal efforts and more “traditional” ways of combining data modalities.

Join the discussion

Changelog++ members save 26 minutes on this episode because they made the ads disappear. Join today!

Sponsors:

  • Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs. 
  • Changelog News – A podcast+newsletter combo that’s brief, entertaining & always on-point. Subscribe today. 

Featuring:

Show Notes:

Something missing or broken? PRs welcome!

More from Practical AI

All 157 episodes
Udio & the age of multi-modal AIPractical AI · 39 min
Listen in VO