In short
How Anish Atraya uses AI models to create AI-generated music videos (Tiny Desk-style), then applies multimodal AI to extract metadata from personal media (record collections and books), plus a quick mention of browser automation for personal finance.
Guests
Anish Atraya, general partner at Andreessen Horowitz and AI consumer investor; long-time DJ/music creator (30 years), interested in remix culture and audio/video generation.
Key claims
AI enables “remix culture” at the video level; simple prompts plus the right tools can produce highly specific aesthetics; constraints (short clips) can increase creativity; multimodal video analysis can automate cataloging; AI browser agents can make websites useful for finance.
Notable examples
“Notorious B.I.G.” and “Kurt Cobain” Tiny Desk concerts using GPT-4/4.0 for prompt/image generation, Hydra for frame-to-video + lip sync, audio stem extraction (DMUX), Audition for editing; a Veo 3 mini music video with 1990s Seattle grunge/camcorder look; Gemini Flash app that scans a video of book flipping to extract author/title and deploys via Cloud Run; Comet RPA for Robinhood portfolio questions.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOCreative AI Music Video Creation
0:00 to 0:50
Learn about the process of creating AI-generated music videos inspired by grunge aesthetics.
“It's like the most creative satisfaction I've had in my whole life.”
Creative AI Music Video Creation
2:03 to 3:02
Learn about the process of creating AI-generated music videos inspired by grunge aesthetics.
“With new AI meeting notes, enterprise search, and research mode, everyone on your team gets a note taker, researcher, doc drafter, brainstormer.”
AI in Music: The Intersection of Creativity and Technology
3:06 to 5:32
Discussion about the evolution of music creation and the impact of AI on creativity.
“Anish, I am so excited to have you here.”
Creating a Notorious B.I.G. Tiny Desk Experience
5:35 to 7:39
Exploration of creating a virtual Tiny Desk concert featuring Notorious B.I.G. using AI tools.
“So I definitely think we're seeing this not just the audio side, but also at the video side, which brings us to your use case.”
Technical Workflow for AI Music Video Production
7:42 to 14:00
An in-depth walkthrough of the tools and processes involved in producing AI-generated music videos.
“4.0 is the best general purpose multimodal model, in my opinion.”
Exploring Constraints in AI Music Generation
14:00 to 15:24
Learn how creative constraints can drive innovation in AI-generated music.
“But one of the limitations I know, having used some of these audio and video gen tools, is you're getting small clips right now with what we're working with.”
DMUX Technology for Vocal Extraction
15:24 to 17:26
Discover the DMUX technology that extracts vocals for creative projects.
“My complaints are so ridiculous because the idea of creating something like this even a year ago sounds so, as you said, impossible that we get so spoiled once we get used to these tools.”
Prompting Techniques for AI Generation
17:26 to 19:12
Understand how different prompting strategies influence AI outputs.
“I think you've got to give the AI the space as well.”
Creating AI-Generated Music Videos
19:12 to 22:25
Explore the process of generating music videos using AI tools like VO3.
“He even manages his mic well, you know, pulls back on some of those notes.”
Refining Music Video Aesthetics
22:25 to 25:41
See how to refine prompts and assemble AI-generated clips into cohesive music videos.
“So this is these are all the videos that generated Google Flow.”
Show all 19 chapters
The Future of AI in Creative Projects
25:41 to 28:03
Discuss the potential of AI in various creative endeavors, including music and education.
“and I also like music videos are a lost art form.”
Using Flash for Video Analysis
28:03 to 29:19
Learn how Flash, a video analysis model, can catalog records and books.
“And the model that does this really well, actually, is Flash, Gemini Flash.”
Building an App with AI Studio
29:20 to 30:59
Discover how to create an app using Google AI Studio for cataloging books.
“This feels like somebody just took a blank piece of paper and brought the best manifestation of the Gemini models forward.”
Innovative Uses of Video in AI
31:00 to 33:09
Explore novel ways to leverage video technology for various applications.
“And so I really think it's great that you're coming to this from how can I solve this with audio?”
Transforming Individual Tool Creation
33:10 to 35:38
Understand the shift towards individuals creating personal software tools.
“And so maybe this will just inspire somebody to build their own record collection extractor, which might be faster than trying to find yours online and reusing something somebody built.”
Consumer AI Experiences
35:39 to 37:38
Examine how AI is impacting consumer experiences in daily life, especially for parents.
“I have to call out as we hop into our lightning round, one thing I noticed, which is you are using Comet.”
AI in Education and Social Learning
37:39 to 41:21
Learn how AI can facilitate social-emotional learning in classroom settings.
“What are the consumer side things that you're excited about?”
Overcoming AI Challenges
41:22 to 42:01
Discover effective strategies for dealing with AI shortcomings and prompting techniques.
“You have had such success with generating these complicated assets.”
Embracing Creative Abandonment
42:01 to 42:30
Learn the value of abandoning unproductive creative efforts.
“Just abandon the branch and start over because you didn't actually do any work.”
Transcript
Automatic transcript. May contain errors.0:00Anish Acharya:It's like the most creative satisfaction I've had in my whole life. So I generated all these clips in a pretty straightforward way. I used GPT-40 to help me with the prompts. I said, hey, help me capture grunge 1990s Seattle inspired by some of these music videos. And then as you can see, it gets progressively more like camcorder, grimy. So I generated all this stuff and then I threw it together into a music video. All right, let's watch it.
0:28you get the patented clairvaux raised hands reaction on this one i cannot believe this is ai generated it's so high quality it's so specific in an aesthetic in a wardrobe an emotion you have inspired me after this podcast what music video am i going to make it's so much fun
0:50Welcome back to How I AI. I'm Clara Vo, product leader and AI obsessive here on a mission to help you build better with these new tools. Today, we have a fun and inspiring episode with Anish Atraya, general partner at Andreessen Horowitz and AI consumer investor. But we're not going to talk about portfolio companies or the future of AI. No, we're going to use AI to build music videos, analyze our bookshelf, and help us plan our personal finances. Let's get to it. To celebrate 25 ,000 YouTube followers on How I AI, we're doing a giveaway. You can win a free year to my favorite AI products, including VZero, Replit, Lovable, Bolt, Cursor, and of course, ChatPRD by leaving a rating and review on your favorite podcast app and subscribing to YouTube.
1:42To enter, simply go to howiaipod.com slash giveaway, read the rules and leave us a review and subscribe. Enter by the end of August and we will announce our winners in September. Thanks for listening. This episode is brought to you by Notion. Notion is now your do everything AI tool for work. With new AI meeting notes, enterprise search, and research mode, everyone on your team gets a note taker, researcher, doc drafter, brainstormer. Your new AI team is here, right where your team already works. I've been a longtime Notion user and have been using the new Notion AI features for the last few weeks.
2:27I can't imagine working without them. AI meeting notes are a game changer. The summaries are accurate and extracting action items is super useful. for standups, team meetings, one-on-ones, customer interviews, and yes, podcast prep. Notion's AI meeting notes are now an essential part of my team's workflow. The fastest growing companies like OpenAI, Ramp, Vercel, and Cursor all use Notion to get more done. Try all of Notion's new AI features for free by signing up with your work email at notion.com slash howiai. Anish, I am so excited to have you here. And let me tell you why. It is because I have spent the majority of this podcast talking about enterprise B2B product management, how to manage your manager or manage yourself as a manager, or how to vibe code.
3:22That has been the topic of how I AI. And today we are just going to have a little bit more fun. So why did you start to come to these AI projects that are a little less like work related or technical and actually just a little bit more fun. How did you get here?
3:41Anish Acharya:Great. Well, I'm excited to have some fun today. I mean, I've been passionate about music forever. I think most of us are. I've been DJing and making music for 30 years, but music is very constrained. There's only so many ways you can work with it. An example of that is if you look at a track that has all the instruments mixed down into a final mp3 or wave file there's no way to just extract the vocal or just extract the drums so you're really limited by a set of choices that were made in the studio and with ai you can do all this crazy stuff like disentangle a track into just the vocals and just the instrumentation so what really got me excited at first was everything you could do with ai and audio and then that of course fed into all of the new video models and video gen and lip sync and all the new technologies we're seeing.
4:26Anish Acharya:So it's just, it's like the most creative satisfaction I've had in maybe my whole life. Yeah, I agree with you. One of the things that I have so much fun with AI on is people are really worried that it takes away the most fun, most human, most creative parts of not just building things, but creating music, creating art, creating writing. And I, in fact, feel like it just gives me so much more tools, so much more breadth, so many more things I can play with and build. And And so it really opens up this like creative artist side of me in a way that has been really hard to access as an adult, also with limited time.
5:02Anish Acharya:Yeah, no. And it's actually a fun conversation we'll have over a glass of wine sometime. But if you look at music culture, music culture has kind of been defined by remix culture for the last 40 years. You know, like the mixtape was the first time that you could take the music and do something, you know, the cassette tape and do something of your own with it. And then that, of course, evolved into, you know, hip hop, which also sampled and which also had a lot of suspicion on it. But sampling was the foundation of hip hop. And I think AI is just the next manifestation of sampling and it'll be as important for music as hip hop was.
5:31Well, and we'll stop opining about AI and the arts. But the other thing that this remix culture makes me think about is kind of the next step that we've seen in the past couple of years, which is kind of audio and video remixing this like TikTok memes, these dances, these things where you're taking a snippet of creativity, turning it into your own thing and then releasing it to the world in a new version. So I definitely think we're seeing this not just the audio side, but also at the video side, which brings us to your use case. So tell me what you built or what you created, maybe. And I'm excited to walk through how you got it done.
6:09Anish Acharya:Amazing. Amazing. Great. Tiny Desk is the best. So if you haven't gotten into Tiny Desk, most people have seen it. It's just it's so cool. It's so fun. And of course, you know, like creativity loves constraints and the constraints of Tiny Desk are incredible. There's a really good one from Clips that just dropped last week. And I mean, anyway, there's an infinite number of them. It's a fun format. It's sort of like the unplugged format of the 90s. So I love Tiny Desk and I got to thinking about all the artists I'd want to see on Tiny Desk. And, you know, of course, some of them are no longer able to be on Tiny Desk because they're not alive anymore.
6:45Anish Acharya:So that got me thinking about how I could do a Notorious B.I.G. Christopher Wallace Tiny Desk. And do we have the tools and technologies? And of course, can we do it in a way that's respectful and not derivative? And I did it and it seemed like it kind of worked. Maybe we can cut to it so your audience can check it out. And the workflow is pretty simple. We'll do a little clip of it, I think, and then we can work through how it got there.
7:22Back to the club, sipping my wet is where your mom was. Back to the club, making holes, my crew's behind. A mad question asking, lung passing, music blasting. But I just can't quit because one of these... Okay, we love it. It's great. And you made that.
7:38Anish Acharya:I did make it, yes. And it took surprisingly little time. Yeah, so let me show you exactly how I made it. So I started with 4.0. 4.0 is the best general purpose multimodal model, in my opinion. I use it for everything. and I just ask it to generate an image. We're going to do Kurt Cobain. That'll be fun today from Nirvana, of course. That's from when I was in high school playing a Tiny Desk concert. So let's see what it comes up with. While this is loading, you mentioned that 4.0 is the best kind of multimodal, all-purpose model. I generally agree. You know, 4.0 ImageGen had this super viral moment a couple months ago when they released it.
8:18What do you feel like 4.0 ImageGen is particularly good at compared to some of the other image gen models.
8:26Anish Acharya:It's very good at prompt adherence. So you can do things, and I think that's because of the infrastructure underneath it. It's a different infrastructure from the diffusion-based models that preceded it. And BFL, Flux, a bunch of others do this now as well, and it's great. But I think it was just the most productive image model because you could manipulate it in such a fine-grained way. Yep. And I remember the biggest improvement when the 4.0 Image Gen came out is that it could actually spell things and write letters out. That was a magical moment. So I have to call out that NPR in the top corner of this image is actually done correctly.
9:00Look, there he is with his cardigan.
9:04Anish Acharya:Okay, I'm going to remove the guitar, actually, so that it is acapella, because I think that might work a little bit better. But look, this is the vibe of Tiny Desk. You know, it's as if you're seeing a photo from the 90s in the Tiny Desk studio. So I just I love this. And I think that we become so attuned to what's possible. We forget that this would be, you know, witchcraft three years ago. Witchcraft. Right. What is the purpose of this? Are you storyboarding? Are you creating an asset that's going to go into another tool? Why start with it on this flow? So so I'll talk through essentially what I'm going to do.
9:37Anish Acharya:So there's this product called Hydra, which is the best way to, I think the best way to take a still frame and add custom audio to it. So create a video that has sort of animated from the still frame and includes the audio with the right lip sync. So, and there's a bunch of amazing tools to do this. Sync Labs is one of my absolute favorites as well. But Hydra is nice because it actually generates the video. So it does the frame to video, and then it also adds the audio. So what we're going to essentially do is take this frame. We're going to get the audio from YouTube. We're going to stem separate the audio so we get the audio track we want.
10:17Anish Acharya:And then we're going to put them together in Hydra. And that's it. This really is remix culture. It's amazing, isn't it? It is amazing. Okay, so the asset that you really need to go into this VideoGen lip sync tool are two things. You need a still image that can be used to generate the video. And then you need some sort of audio to sync this to. So I know we're looking at this music example, but what other examples have you seen people use this kind of workflow for? I think we underestimated how useful it would be to add custom audio to video. And there's been a bunch of great, you know, one of the early examples was taking a speech that somebody was giving.
10:57Anish Acharya:I know Javier Millet did a really famous one. And essentially lip syncing, changing the language to English and lip syncing it. That went really viral a couple of years ago. So we've seen. And then, of course, you can imagine a character, a photo of a character that you generate. And then you want to animate them doing something and speaking at the same time. So, you know, stories are told this way. And these technologies make it really, really easy to do so. Oh, we got him. Great. Okay. So now he's got bad posture, but we'll allow it. It's very grudges about it. I think he always did. Yeah. Exactly.
11:30Anish Acharya:Okay. So now we've got Kurt. Now, what I would do if I didn't actually have a... So Tiny Desk has got a really specific acoustic aesthetic, which is it sounds like live instrumentation. So for the Biggie example, I actually found a Biggie cover band playing live in Brooklyn. And I pulled that down from YouTube. And then I extracted the actual vocals from the Notorious B.I.G. and laid them over. But in this case, Nirvana did a really famous New York City Unplugged concert in 93. So there's video of them playing in the way that they would and audio in the way that they would on Tiny Desk. So that is right here.
12:10Even in the same cardigan.
12:12Anish Acharya:Even in the same cardigan. Isn't that amazing? Yep. Okay. So I use this nifty little tool called 4K Video Downloader, which is slightly sketchy, but that's okay. I love these little utilities that you just, you know, you Google like, how do I get audio out of YouTube? And then you look at the scariest website possible. And you just cross your fingers that your computer won't go up in flames and you download 4K video. Yes, my yes, my data is definitely going somewhere sketchy as a result of this. So for the vibe coders that are listening, I have a request for startup, which is go find all these slightly scary little utils and build me ones that are less sketchy looking.
12:56Anish Acharya:100%. 100%. It's a great idea. Okay. So now we actually have this. So we've got the video. Yep. Now we're going to open Adobe Edition. Okay. So this is a tool that people who have been working in computer audio have been using for 30 years plus. It used to be called CoolEdit Pro. It's completely beloved. and it's very, very easy to use, which is why so many of us use it. It was, of course, acquired by Adobe many years ago. It's now called Audition. So I go to Audition, and I take this video, and I just drop it in. So here we actually have the audio from the video, which is really, really cool.
13:33Anish Acharya:I'm going to zoom in, and I'm going to see the first few seconds of it are blank. So let's just cut that out because we don't want to hear that. then we're going to zoom out and we're going to take, I don't know, let's take 15 seconds. And you can kind of see the audio, the video in the bottom left corner there. Oh, got it. So it's combining the audio and video just so you know exactly what you're syncing up to. Exactly. And I'm going to pretend like you're doing 15 seconds because we're doing a very efficient podcast here. But one of the limitations I know, having used some of these audio and video gen tools, is you're getting small clips right now with what we're working with.
14:13And so, you know, what I'm looking forward to is the day where I can have the, you know, hour-long Nirvana unplugged tiny desk. Totally. But, you know, do you ever feel constrained by the kind of length of assets being generated or the quality?
14:30Anish Acharya:I mean, sort of. But again, I think creativity breeds constraints. So not to over-rotate on hip-hop, but if you look at the reason that so many samples were used in hip-hop in creative ways in the 80s and 90s was the actual drum machines and samplers had very limited sampling time. So you could only sample a second of anything, so you couldn't really sample four bars. And that's why so many producers put tracks together that use these many one-second samples in surprising ways. And once we actually got the technology to sample for more time, we actually got less creativity, I would argue. So I sort of love the constraints that the technology gives us today.
15:08Well, I also love my complaints. I'm like, isn't it annoying that you can't provide Nirvana and overlay their audio and generate a completely fictional concert for longer than 15 seconds in probably under a 30 minute podcast? My complaints are so ridiculous because the idea of creating something like this even a year ago sounds so, as you said, impossible that we get so spoiled once we get used to these tools.
15:36Anish Acharya:100 % right. No, exactly. I mean, this stuff, we would have called it witchcraft three years ago. It would have been. Okay, now there's two things you can do with this. If we wanted to do an acapella-only version, for example, we can use a technology called DMUX. So DMUX is this amazing technology that allows you to extract the vocals from any song. So here, I've forgotten what the actual command line is. So I just do this. I looked it up in perplexity. What's the actual way to extract two tracks with DMUX? We do this, DMUX two stems vocals, and then let's go find the path. Okay, so this command is going to take that audio file we saved of the first 15 seconds of this concert, and it's going to extract the vocals from the instrumentation.
16:28Anish Acharya:So this will be Kurt Cobain singing come as you are acapella which as far as I know has never happened which is pretty cool and then we simply come back here and we say start frame upload an image let's use this okay that's our Kurt Cobain audio script upload audio and let's use actually the full audio with all the instruments add a video and then we just say man singing on tiny desk what i love about your prompting compared to other how i ai guests is every prompt has been sub six words six words you're very simple in terms of describing what you want and uh get high quality outputs there so i don't know what that says about the prompt engineering industrial real complex.
17:20But prove here that you can use simple prompts to get pretty cool stuff if the tool behind the scenes does the work for you.
17:27Anish Acharya:I think you've got to give the AI the space as well. You know, if you overly constrain it, it just really struggles to satisfy you. Whereas if you give it less constraints, you know, sometimes it has unexpected results, but often they're unexpected, you know, delightful. Well, that's what I've heard a lot from folks that come from the more creative backgrounds. Designers in particular tend to be less precise in their prompting because they want that exploration space that then they can narrow in on. And so I really think it also comes into play, your prompting technique can come into play based on kind of what profession or what background you're coming from.
18:04Engineers want like the most precise. They not only want the code to work, but they want the code to be written exactly how they would write it. And so they're very precise in their prompting. Where I found designers and more creative folks building different kinds of assets really like that wide open space.
18:20Anish Acharya:Totally. Yes, exactly. And while we're waiting for this to load, it might be interesting. I'm just looking at some of the options at the bottom here. So you have different kind of models that you can use, including one that looks like that they specifically fine-tuned for this, different aspect ratio, orientation length probably based on the script and then you know the prompt says prompt your character with emotion and gesture so i am very curious if you put like angsty man singing versus cheerful man singing if you'd get a different a different version here even if the audio video were were the same it works really well absolutely yeah no this this is such a useful storytelling product it's it's amazing and when you combine it with other video gen models like vo3 you can start to tell real stories, you know?
19:11Anish Acharya:Yeah. Okay, let's check it out. All right.
19:24All right.
19:25Anish Acharya:Pretty cool. It's very good. It's very good. Very satisfying. He even manages his mic well, you know, pulls back on some of those notes. Totally. That's incredible. And so, you know, could you take this and take different clips of the video and sort of generate a string of these these videos and maybe put them together in a longer form version? A hundred percent. Yeah, I actually was inspired by this. So I put together a music video, a little mini music video for a different Nirvana track. Can I show it to you right now? Yes, we would love to see it. Okay. I used VO3 to generate the clips and And it turned out great, I think.
20:09Anish Acharya:Hold on one moment. Yeah, and I think if you haven't tried VO3, it is pretty incredible. I mean, I can only generate like two and a half videos every day of three, you know, seven second length or whatever. I'm still capped on usage, but the quality is really good. The physics are really good. It's one of my favorite video models to play with right now. Just as a consumer, it's kind of, to me, my experience with that model was very similar to my first experience with Mid Journey, where just the breadth of things coming out of the model were so incredible to me. So highly recommend folks give that model a little spin.
20:54Anish Acharya:It's amazing. Yeah. You've got to get on Gemini Ultra, Claire. We have a household Gemini Ultra. account. But my husband is the video gen guy. So he's up there. And by the time I get to it, we burn through some tokens. But I spent all the money on Cursor. Fair. I know. My wife, for the first time this month, was like, babe, what is Cursor? I'm like, ugh, don't worry about it. I know all these little secret AI tools popping up on the credit card. How I AI is now on Lenny's list with my personal selection of the best AI engineering courses on Maven. You can spend months thinking and playing with AI before really integrating it into your workflow or shipping an actual AI feature.
Read the full transcript
21:51If you want to start building, then these hands-on Maven courses are for you. Learn directly from Aishwarya Naresh Riganti, MIT instructor and AI scientist at AWS, or Sandra Shuloff, who has authored research with OpenAI, Hugging Face, and Stanford. To pivot into an AI role or successfully lead your company's next AI initiative, visit maven.com slash Lenny to enroll now. Use code LENNYSLIST for$100 off. That's M-A-V-E-N dot com slash Lenny to get ahead in the AI era and start building.
22:35Anish Acharya:So this is these are all the videos that generated Google Flow. So I was trying to capture like a 1990s high school band auditorium, you know, a little dystopian energy. So I generated all these clips in a pretty straightforward way. I used GPT-40 to help me with the prompts. Because as you can see, this is actually the beginning of my generations. This is like the complete wrong energy. I don't know what this is, like early 80s synth pop or something. So then I went to GPT-40 and said, hey, help me capture grunge 1990s Seattle, inspired by some of these music videos. And then as you can see, it gets progressively more like camcorder and sort of grimy.
23:17Anish Acharya:so I generated all this stuff and then I threw it together into a music video and I put the music behind it I'll show it to you right now amazing so just restating this 4-0 helping you refine your prompts to get the aesthetic right the phrasing the prompting right give you some keywords Veo to generate these like shorter clips and then do you put it together in like final cut or something like that I put it together in cap wing cap wing is so easy and so useful. I highly recommend it. I'm a tape top girl, so I use cap cut. Yeah, gotta get on cap bling. All right, let's watch it.
24:20To
24:56that's it okay you get the patented clairvaux raised hands on this one i'm gonna tell you the real truth something like this makes me almost want to cry because i really got in technology i wanted like everybody i want to like make video games and like make movies and work for Pixar or directly. And it always felt so inaccessible to get these like amazing ideas that I had in my head. Into a thing like, could you film it? Could you access the people? Did you have the time? Did you have the music? Did you have to create? And you just put together this amazing, amazing music video. Thank you. I'm so impressed.
25:38Anish Acharya:Thank you. It was so fun. It was so easy. and I also like music videos are a lost art form. Totally. I'm so excited to see, you know, everybody making music videos for all their favorite tracks because what a cool way to contribute, you know, and in no way does it actually dilute from the original. I think it's a testament to the original and our appreciation of it. No, it looks like a love letter. And I have to call out when I was watching it, there's a lot of it that I think is incredible. I like how the cameras, you know, like pan and zoom in. the part that really got me was the sequential shots of the teenagers in the hall and i was like i cannot believe this is ai generated it's so high quality it's so specific in an aesthetic in a wardrobe in a motion and it got me until and again they are three good at physics until there's like a guy with like a pack of camel cigarettes on his arm and like the cigarettes So like halfway coming out.
26:34Anish Acharya:Yes, yes, yes. Totally. That's right. Well, actually, and the other funny artifact is if you look at the end when the band is flying and a bunch of people are jumping out of the crowd, four people jump out of the crowd at the same time. They look the same and they're making the exact same like, you know, they look like acrobats at a circus or something. It's like the end of like an 80s TV special where they all jump up. Totally. Yes, yes. That's amazing. You have inspired me truly after this podcast. I'm like, what music video am I going to make? It's so much fun. Do it. Do it. I mean, music videos.
27:09You could do like fake movie trailers. Yes. Also documentaries. I mean, we're doing the fun art, you know, heart and soul filling stuff. But I also think the ability to create educational materials that are compelling and interesting with this technology are also right there.
27:28Anish Acharya:I mean, if you look at fan fiction, fan fiction's enormous because people want to contribute to the things they love. And now we get fan fiction for every medium. It's so cool. Okay. Sold. All right. That was just workflow number one. We're going to go pretty fast through workflow number two, which I think is a little bit more of a practical one, but still connected to the arts. So walk us through what your second workflow is. Cool. Yeah. So one of the things that I think is really under hyped, underappreciated, underused is all of the multimodal capabilities. And the model that does this really well, actually, is Flash, Gemini Flash.
28:08Anish Acharya:So it's just it's great. It's one of the very few models that can do video analysis and ingestion. It can do all kinds of amazing things. And yet I don't see it being used out there a lot. I thought I would use it to create an app that would help me catalog my record collection. because I've got, you know, like every DJ, I've got so many records and it's such a pain to keep track of them and know which ones I had and which ones I didn't. So I did a very quick app on Friday that let me take a video of flipping through my record collection and then using Gemini to extract artist names, album names, photos.
28:41Anish Acharya:It's really, really cool. So I thought today we could do something similar except for books. This is amazing. And we were talking before we started recording, this is going to help me because over here, I have like 100 books and 100 records piled up on shelves that have definitely not been cataloged. So I can't wait to see what this looks like. Perfect. I got you. Let's share. So here we are in Google AI Studio. So I'm sure folks are familiar with AI Studio. But if you're not actually, I think it's the best product surface to interact with all the Gemini models, one of the best anyway, because it doesn't have all of the kind of overhead and links and constraints that a lot of the other Gemini products have.
29:24Anish Acharya:This feels like somebody just took a blank piece of paper and brought the best manifestation of the Gemini models forward. So I really love AI Studio. That's my starting point for all of these things. And then in AI Studio, you can see here, you can of course chat, you can stream with your phone or with your webcam, you can generate media, and you can build apps. This is a very good app builder. And this is the best way to build off the shelf apps, I think, that integrate with Google models. So here I've typed, you know, create an app that takes a video of a person flipping through their book collection and extracts the author and title of every book shown.
30:01Anish Acharya:Then I give it a suggestion for how it could do it, which is you could do this by taking the video and first extracting the frames that show distinct books and then have a vision model analyze those frames to extract the information. Make sure you extract every book shown sequentially. What I have to call it here is, you know, what's interesting is people know that these models exist and they generally know some of the capabilities, vision, you know, text to speech or speech to text, all this stuff. But what's really hard for people to do, and I appreciate you showing us is think of novel ways you can access the abilities of those models.
30:42I would have actually, I thought you were going to show us like you took a picture of it and you cataloged it. But this idea of a video and then extracting the frames, I just haven't changed my mental model to match these multimodal models in order to take, you know, take advantage of things that can be more efficient, allow you to do things. And so I really think it's great that you're coming to this from how can I solve this with audio? How can I solve this with video? How can I solve this with text? And knowing that the models can do kind of the hard work on the back end.
31:14Anish Acharya:Thanks. Yeah, look, I completely agree. And video is just, of course, it's so much more rich than image. And this is the way that we bring a lot of the outside world online, I think. So I've been really inspired by video. I saw something on Twitter where somebody had set up a mini app that watched him shoot free throws and kept count. You know, you could, I mean, there's just so many ways that this will be productive. I'm very passionate about AI for parents, and I've got kind of a neat video idea in there as well. So to me, there's like the sort of skeuomorphic technologies, which is using the new technology with the old assumptions.
31:47Anish Acharya:And then there's the native ways to use it. And this feels like a very native way to use the models. Well, to connect the two things that you said, the, you know, basketball shooting analysis in kids, my husband did upload every single one of our eight-year-old basketball games to a video analysis to get like each each kid's step no way shooting percentages all the they actually don't even keep score at this age so he got to like get the score i love that yeah i totally love okay so now we have an app yeah so i'm going to take a video here of me just flipping through my stack of
32:30books.
32:31Anish Acharya:Okay. I've taken the video. Okay. And that took all of seven seconds. So yeah, exactly. Yeah. Now, you know, the one edge here that's kind of interesting is this is really, it's really easy to get something working, but if you want to publish an app that a lot of other people can use, it then becomes more work. So I probably, it took me 15 minutes to create this for my record collection, at least create the working demo in Primitive, but then it took me half a day to get it live so anybody could use it. And what's interesting about that is I feel like a lot of individuals are just going to build their own tools and presume other people are going to build their own tools.
33:13And so maybe this will just inspire somebody to build their own record collection extractor, which might be faster than trying to find yours online and reusing something somebody built.
33:22Anish Acharya:I mean, the era of personal software is upon us, you know? Totally. Okay, so what it's doing, taking this video, it's going to do frame-by-frame extraction of the, again, something that is just so time-consuming. And then it's going to use the vision capabilities. What model do you know is behind the scenes of all this? You say Flash? It's Flash, Flash 1.5. And I can kind of skip ahead and show you what this looks like. So here is, here's one that I built yesterday and which with essentially the exact same prompt. Yep. So let's run it in parallel and see if this one's any happier with us. Okay.
33:58And I did notice one was light mode and one was dark mode. Was that just the AI god?
34:03Anish Acharya:Yeah, this is just some of the randomness of the models. Yeah, exactly. Oh, I do, I do have to say I like the progress indicator of the second one. It told me how many frames it's extracting. Oh, look at this. So here we go. You know, this is the Chris Dixon book, the Paul Graham book, this very nerdy book that Mark asked me to read when I was hired. This is a really good Thomas Sowell history book. Anyways, this is my entire stack of books, every single one of them. You can see a photo. It's extracted the author and the book name. so it's like you know this is just a couple of prompts that's it and it generated it so this is what's possible and then if you go here to deploy with cloud run you get a deployed version of it that's actually running on the cloud and now you can send this link to anyone now this is going to cost you api credit so maybe you want to be a little bit deliberate but you're pretty much ready to go with this really sophisticated video processing app that would have taken i don't know, a month of time previously.
35:06Yeah. Amazing. And so useful because now I can figure out which of these also very nerdy books we have. We've read. I also see some duplicates up there. Totally. Yes. Yeah. Not perfect. Yeah, exactly.
35:20Anish Acharya:Well, actually, in this case, the photo is duplicative, but it detected the Ben book and the Chris book separately. So, but yes. I need this, man. I need this for the pile of kids books I have up in my kids' closet so they even remember what they have. Okay, this is great. Well, thank you so much for showing us these fun use cases. I have to call out as we hop into our lightning round, one thing I noticed, which is you are using Comet. I am using Comet. Tell me a little bit more about why that new browser is your browser of choice and what are you getting out of it? Comet is so good. I mean, I've been skeptical of the new browser thing because it just feels like the ways to improve the browser in the past have been very incremental.
36:07Anish Acharya:Ambitious, but there just wasn't that much surface area for new browser features. And now with Comet from Perplexity, it can do a bunch of really incredible things. My favorite thing that it can do is what's called RPA, which is where the models operate your browser on your behalf. So you've seen a bunch of examples of this, of like, hey, go find me a flight and pay for it, which is interesting. the way I've been using it is in my finances. So I'll go into Robinhood and I'll say, hey, why don't you tell me how my portfolio is performing? Why don't you tell me where I could get stocks that have similar upside at a lower cost basis?
36:43Anish Acharya:What stock should I buy next? Are any of these memes? I mean, you can just go so deep. And look, I could probably figure that out by clicking around the website and downloading the data, but now I don't have to. So this assistant feature in Comet makes every website dramatically more useful. And it's been a big unlock for me. I love this whole episode because you've actually shown a couple of use cases, including talking about personal finances with Comet that really are consumer use cases. Again, as I started at the beginning, we're doing a lot of like, how do you work this inside of an enterprise?
37:19How do you write code with it? But I think the real underappreciated transformation is gonna come in consumer experience. I think we're so early. I mean, as somebody who does a podcast trying to educate people, I just realized we're so early on consumer adoption of AI. And so I have a question for you, which is if you could get, you know, like my mom or, you know, one of my friends that is less, you know, not in Silicon Valley, less in the middle of this in a room and say, you know, let me show you three things in 15 minutes that are just going to totally change how you think about your life or things that you never knew were possible, what would be those things?
37:58What are the consumer side things that you're excited about?
38:01Anish Acharya:So I have kids and parenting is on my mind all the time. And the ways that my kids use models are amazing. So for my four-year-old, Chachi BT reads her a bedtime story, but not just a bedtime story, one where she can ask infinite questions. So what was the king's dragon's name? What color was it? Where did it come from? Did it have any kids? You know, she's really into unicorns and alicorns, like tell me a story about an alicorn and a golden egg. And so she can just really interact with the bedtime story. And ChatGPT is far more patient and creative than we usually are. So that's one way. And look, she can't really use a computer otherwise, other than watching YouTube.
38:42Anish Acharya:And then for my son, he'll set up two figures like Sandman and Spider-Man, and then he'll take a photo of them in Chachibitir, one of the other models, and say, hey, who would win? And then it'll do this whole, you know, oh, Sandman would win in these conditions, and Spider-Man, but maybe Spider-Man does this. So they're able to kind of play with the technology instead of just being broadcast to from technology, which is really new. That's like the near-term stuff. I think in the longer term, you know, I think that the models can really help with a lot of social-emotional learning. If you look at the classroom, Part of it, of course, is academics, but part of it is just teaching children to be, you know, good people for the world.
39:25Anish Acharya:And a lot of that comes in observing how they're sort of behaving and interacting. And we never had a technology that could do that. If your kid went to a great school, there might be a second teacher in the classroom focused on social emotional. So I think that's how AI shows up in the classroom. It's probably less like homework helpers and assignment generation and more observing the social dynamics in a classroom and helping kids be better people. Yeah, well, calling back to what we were saying earlier about trying to identify the AI native way of doing things. I watch my children so much. I say that my children form my consumer AI theses for me.
40:03Because the other day, my six-year-old was playing Minecraft and he wanted to know how to do a command. And he literally went to my purse, picked up my Meta AI glasses, put them on and said, Hey, Meta, how do I transport to the Woodland Mansion in Minecraft? And I was like, wait, this is like, it's not type into chat GPT. It's not even ask Alexa. He took this physical device and put it on his face and asked this personal AI a question. And that just really opened my mind to, again, I think multimodal is going to change. I think hardware is going to have a real place to play here. And then this like AI native generation is going to think about accessing information and building things in a totally, totally different, totally different way.
40:52So I am with you on all of that.
40:56Anish Acharya:I love that. Yeah. And it's interesting because we have been taught what computers can and can't do, but they haven't been taught any of those things. So when I generate an image of, you know, a Harry Potter image for my son, I'm like, wow, do you see how I just generated that? He's like, dad, of course the computer can do that. So they just assume that everything's possible. And now everything kind of is. Oh, my gosh. We had it. As I say, when I had to walk uphill both ways for my Internet. That's right. You and me both. We'll get you out of here. One last question I have to ask. You have had such success with generating these complicated assets.
41:30But when AI is not listening to you, when it is giving you really poor results, what is your prompting technique to get it back on track?
41:38Anish Acharya:I mean, I don't know if it's a prompting technique, but it's a mindset. Two things. One is go with it. Let it take you to some strange, unexpected places, and you might be amazed at the results. I think the other is just reducing this sunk cost fallacy thing where you create a GitHub branch, you try to do something really ambitious. It's just like falling over, over and over again. Just abandon the branch and start over because you didn't actually do any work. you feel like you did work because it did work, but that's not you doing work. And I think being a lot more willing to abandon sort of approaches that aren't working is the sweet spot.
42:17I completely agree. Well, thank you so much for showing us all these works flows. It was totally inspiring. I want to get off this podcast so I can go play. So thank you for making my day. And I know everybody's going to love the episode.
42:30Anish Acharya:Thank you, Claire. Super fun. Thanks so much for watching. If you enjoyed the show, please like and subscribe here on YouTube, or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiaipod.com. See you next time.
From the publisher
Anish Acharya is an entrepreneur and general partner at Andreessen Horowitz, focusing on consumer investing and AI-native products. In this episode, he demonstrates how AI can be used for creative and personal projects beyond typical work applications. He walks through creating an AI-generated Tiny Desk Concert for Notorious B.I.G. and Kurt Cobain, building a book cataloging app using video analysis, and using browser automation for personal finance insights. Anish shares how these technologies allow anyone to bring creative ideas to life with minimal technical expertise, transforming what would have been impossible projects just a few years ago into accessible weekend activities.
What you’ll learn:
1. A step-by-step workflow for creating AI-generated music videos featuring artists like Kurt Cobain and Notorious B.I.G.
2. How to extract vocals from existing tracks to create unique audio combinations for your AI-generated videos
3. A simple method for cataloging your book or record collection using video analysis and Gemini Flash
4. How to use Comet to analyze personal finances and get investment recommendations without manual data analysis
5. Ways AI is transforming childhood learning and play by enabling interactive storytelling and creative exploration
—
Brought to you by:
Notion—The best AI tools for work
Lenny’s List on Maven—Hands-on AI education curated by Lenny and Claire
—
Where to find Anish Acharya:
• Andreessen Horowitz: https://a16z.com/author/anish-acharya/
• LinkedIn: https://www.linkedin.com/in/anishacharya/
—
Where to find Claire Vo:
ChatPRD: https://www.chatprd.ai/
Website: https://clairevo.com/
LinkedIn: https://www.linkedin.com/in/clairevo/
—
In this episode, we cover:
(00:00) Introduction to Anish Acharya
(03:05) How AI transforms creative constraints in music and video
(06:00) Creating an AI-generated Notorious B.I.G. Tiny Desk Concert
(07:36) Using GPT-4o to generate still images
(09:27) Using Hedra to animate still frame images
(10:40) Adding custom audio to video
(11:30) Using Adobe Audition to clip and sync audio
(15:42) How to use Demucs to extract vocals from any song
(16:36) Using Hedra to generate a Tiny Desk Concert featuring Kurt Cobain
(19:40) Creating a ’90s-style Nirvana music video with Veo 3
(27:40) Building a book collection cataloging tool with Gemini Flash
(35:35) Using the Comet browser for personal finance analysis
(37:20) How AI is transforming childhood learning and play
(41:23) Tips for getting better results from AI tools
—
Tools referenced:
• GPT-4o: https://openai.com/index/hello-gpt-4o/
• Hedra: https://www.hedra.com/
• Adobe Audition: https://www.adobe.com/products/audition.html
• Demucs: https://github.com/facebookresearch/demucs
• Perplexity: https://www.perplexity.ai/
• Veo 3: https://deepmind.google/models/veo/
• Kapwing: https://www.kapwing.com/
• Cursor: https://cursor.com/
• Google AI Studio: https://makersuite.google.com/
• Gemini Flash: https://ai.google.dev/gemini-api
• Comet: https://www.perplexity.ai/comet
—
Other references:
• Anish’s Notorious B.I.G. AI-generated Tiny Desk Concert: https://x.com/illscience/status/1935721063876550939
• NPR Tiny Desk Concerts: https://www.npr.org/series/tiny-desk-concerts/
• Notorious B.I.G.: https://en.wikipedia.org/wiki/The_Notorious_B.I.G.
• Kurt Cobain: https://www.kurtcobain.com/
• Robinhood: https://robinhood.com
—
Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.




