#224 Matt Panousis: How AI Is Disrupting The Dubbing Industry (LipDub AI)

8 Dec 2024 · 52 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. Podcast Episode #224: Matt Panousis - How AI Is Disrupting The Dubbing Industry (LipDub AI)

Episode Overview In this episode of Eye On A.I., Craig S. Smith interviews Matt Panousis, Co-Founder and COO of LipDub AI, a revolutionary AI-powered lip-syncing tool aimed at transforming the dubbing industry. They discuss the evolution of LipDub, content localization, and the challenges and implications of AI in video content.

Key Themes and Discussions

  1. Introduction & Background
  2. Matt Panousis introduces himself and his background as a lawyer turned entrepreneur, highlighting his experience in starting a software company and transitioning into the visual effects industry.
  1. Origins of LipDub
  2. LipDub was born out of the founding team’s experiences in visual effects and recognition of Hollywood's demand for faster, more cost-effective solutions.
  3. Initial focus on Vanity AI, aimed at automating digital makeup and de-aging, which later led to the development of LipDub for lip-syncing dubbed audio.
  1. Technological Foundations of LipDub
  2. LipDub utilizes advanced AI to automatically sync lip movements with dubbed audio, addressing a significant gap in the market.
  3. Discusses the technical process, including:
  4. Detection and tracking of faces in uploaded media.
  5. Training on speakers for improved texture and realism.
  6. Importance of mouth internals and articulation accuracy.
  1. Differentiation from Competitors
  2. Compared to competitors like RasQ AI, LipDub focuses on high-resolution, dynamic content, allowing for complex scenes with multiple characters.
  3. Emphasizes quality and creative freedom as key differentiators.
  1. Market Application and Target Audiences
  2. Current target markets include:
  3. Hollywood for film and TV dubbing.
  4. YouTube creators looking to localize content.
  5. Online education platforms and advertising agencies.
  6. Mr. Beast mentioned as an example of a YouTuber exploring dubbing to increase international audience reach.
  1. Challenges in Dubbing and Localization
  2. Timing and Translation Accuracy: Discusses the challenge of matching lip movements with audio while maintaining the original script's intent and timing.
  3. Need for human oversight to ensure translation accuracy and cultural relevance.
  1. Ethical Considerations and Misuse
  2. Addressing potential misuse of technology, including creating misleading content.
  3. LipDub requires users to confirm they have permission to dub individuals and employs spot-checking for compliance.
  4. Discusses the importance of collaborating with platforms for content labeling to prevent misinformation.
  1. Future Developments and Vision
  2. Exploration of real-time dubbing and its implications for global communication.
  3. Potential expansions into digital watermarking to prevent misuse and ensure authenticity.
  1. Team Building and Research
  2. Insights into how Matt built a strong team, emphasizing the importance of world-class researchers and advisors in AI and graphics.
  3. The role of Daniel Cohen-Or, a leading researcher in the field, in developing LipDub's technology.
  1. Conclusion
  2. LipDub is positioned to redefine content localization and dubbing in the digital age, empowering creators to reach global audiences with high-quality, engaging content.

Key Takeaways

  • AI is poised to significantly disrupt the dubbing industry, making it accessible and affordable.
  • Emphasis on quality and dynamic content will differentiate LipDub from existing solutions.
  • Ethical considerations around the technology's use will be critical for its acceptance and success.

Stay Connected

  • Follow Craig S. Smith on [Twitter](https://twitter.com/craigss).
  • Follow Eye on A.I. on [Twitter](https://twitter.com/EyeOn_AI).

This episode provides valuable insights into the intersection of AI technology and media, illustrating the rapid evolution of tools that enhance content accessibility and engagement in a global context.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Anything that's dubbed deserves to be lip synced. So in terms of circling back to your original question, who's the user? It's a constant conversation. Really excited by the YouTube market. I haven't seen any one particular trend stick out. It just feels like everyone wants to better communicate with everyone. Very few things in Hollywood are just single identities looking at a screen. It's people running away from burning buildings and scenes where you have 10 characters speaking and people are turning their heads to the side and lighting is changing. And that's really where we invested a lot of our R &D work was not only being able to do great articulation and high resolution, high fidelity textures, but also being able to do difficult content or we call it internally like dynamic content.

0:49Okay. So Matt, why don't you start by introducing yourself? Tell us a little bit about your background and how you got to LipDub. Sure. So my name is Matt Nusis and I'm the co-founder of Monsters, Aliens, Robots, Zombies. Lawyer by training. Started my first software company out of law school. So I never practiced. That software company had nothing to do with the work that we're doing here today. It was an e-learning company called Acto. spent the better part of spent the better part of five years working on acto exited that and then uh two partners of my current partners were starting a visual effects company which i was intrigued by i knew nothing about the space how did we get to lip dub And phase one was we were working on this visual effects company and we were seeing this demand from Hollywood for faster and cheaper visual effects.

2:01And so we started to ask ourselves, what would it take to do the effects work or visual effects work at a fundamentally faster pace at a much better price point and without sacrificing quality. And we recognized pretty early, this was back in 2018 when we started that, if we were going to change or have a truly differentiated offering for Hollywood, we would need to invest in innovation. And AI at the time seemed like the right innovation to hang our hat on, given those goals. After making that decision, it was us looking for applications because visual effects, Craig, is very broad. You do a lot of different things when you're working with Hollywood.

2:47You might be doing a creature for Stranger Things, or you might be making the Toronto skyline look like the New York skyline, or you might be de-aging someone, or you might be creating a wave simulation. So it's such a broad, it's such an umbrella term, visual effects, that if you're going to try and innovate in the space, it's really important that you actually choose an application that you're trying to automate. And we were looking for broad applications because obviously the R &D work that goes into building AI products for Hollywood is really intensive and you certainly don't want to spend years building and automating a use case that barely shows up or shows up on one of every 10 projects.

3:34You're really looking for use cases that show up on the vast majority of projects. And the first application that we decided to invest in was called Vanity AI. The reason we liked Vanity was digital makeup work and aging does show up on the vast majority of Hollywood projects. And at the same time, we felt looking at the trends in AI, obviously this predates the gen AI craze that we're in today, but we felt like the tech was getting good enough to accommodate that use case. And so that's where our AI endeavors began. We built an internal tool that we use here at Mars that truncates the time that the FX artist needs to do a digital makeup and or de-aging shot.

4:28So previously, a five second shot might take a the effects artist anywhere from a half a day to depending on what's being asked by the client and how much you're changing the face it could take two three days of artist time so vanity truncated that down to 25 minutes a shot on average so that was our first endeavor and basically we were We're already working on the face. So I think a lot of our research was focused on the face. Our chief scientists, Daniel Cohen-Oar and his lab do a lot of work, what they call deep facial editing. So we were already focusing on that region. And then it just so happened that Squid Game came out shortly thereafter.

5:18Most of us will watch the series. It was a fantastic story, but obviously the lack of synchronization between the lips and the audio took us out of the experience. And we said that could be a really interesting problem to solve. Like no one's been able to solve that to date. This sounds like it's right in the pocket of what we're out there to do, which is to create these kind of highly automated visual effects applications. And so that was the original thesis behind LipDub was let's make Hollywood dubs look real for the first time by automatically syncing the lips to whatever new dubbed audio track is fed into the system.

6:06That was the ethos of where we started. Obviously, now that we're in market, we've learned a lot about other industries and their use cases for LypDub. But in terms of how we got there, that was the evolution. Yeah. And when you say it's a problem that needed to be solved, there are other technologies, other solutions out there. I'm thinking of RASC AI, but they're not as precise from what I understand as LypDub. and for Hollywood, you need a more precise solution. Is that right? Is that what differentiates you guys? Yeah, so for us, it was always about two things. With Hollywood, obviously the quality bar that you're solving for is as high as it gets.

6:55And so a lot of things need to work at a certain level in order to be usable. So obviously the articulation has to be perfect. The fidelity of textures and the resolution that you operate in have to be Hollywood grade, which is typically now 4K and fidelity of textures has to be fantastic. If you have a beard, if we were lip dubbing you, which we're going to do here, we want to be able to see the individual strands of hair on your beard. So that was an important requirement for us in solving for this was quality of articulation, textures and fidelity of textures on the face. The other obvious one is the fact that very few things in Hollywood are just single identities looking at a screen.

7:45It's people running through, running away from burning buildings and scenes where you have ten characters speaking and people are turning their heads to the side and lighting is changing. And so that's really where we invested a lot of our R &D work was not only being able to do great articulation and high resolution, high fidelity textures, but also being able to do difficult content or we call it internally like dynamic content. And so when you think about other tools in the market, like Rask as an example, interestingly enough, so those tools started on the audio side. So their original purpose for being was to automate the dubbing aspect of the equation, which we never worried about because again, if you think about why we solved this problem, it was under the guise that Hollywood would be providing us the audio track.

8:44So that was never a focus of ours. Now some of those audio companies have started to endeavor into lip sync to have the quote unquote, you know, soup to nuts localization solution. But where we differentiate is exactly along the lines that I've just spoke about. So when you work with Lyftup, what you're getting is you're getting the best articulation on the market, the best resolution on the market, and you're not limited creatively. You can do anything with Lyftup. You can do people moving, you can do people speaking in side pose, you can do object interference, things passing the face, and I think those are the big ones, honestly.

9:29Just whatever your video content entails, you're not limited. Whereas most of these more consumer-grade systems tend to struggle with even the basics. Yeah. And then how does your solution or your platform integrate with existing dubbing solutions? obviously 11 labs is the leader i think right now or deep dub another company that does dubbing and you talked about these soup to nuts solutions for more consumer grade are you going to add the dubbing portion to your platform so that people can do it all on your platform? Yeah, it's certainly something that we talk about frequently. So today, most of our clients outside of Hollywood and clients entail advertisers, folks doing either online education for their own employees or let's say selling courses online, YouTube channels, advertising agencies, what we realized is a lot of these users do also need the audio solved.

10:57Advertising is somewhat of an exception. They still do leverage real voice acting, but you can see they're actually starting to shift and move towards these really economical solutions. Our approach today with our current clients has been go and purchase a deep go and purchase in 11 labs and then use us. I think where we're going to move in the future is we're going to be likely wrapping around a tool. We haven't decided exactly which one yet, but we've had a lot of clients asking for a one-stop shop, not to say using two pieces of software is, it's not a non-starter, but having everything in one place they're telling us would be advantageous.

11:40And so that's definitely a conversation that we're having now internally. Yeah. And what's the process, the technical process, the algorithms used to match lip movements or manipulate pixels in the video so that lip movements match the audio, the dubbed audio? Like how does the product work itself? Yeah. Without exposing too much, because a lot of what we do and what makes us unique is the secret stocks and the proprietary work that's gone in over two years with Daniel and his team. Pretty much the way the product works is it operates like Dropbox. So you have an original piece of media. Let's say that media was done in English and you're looking to target Mandarin.

12:31That's exactly what we're going to be doing with this podcast. The process is really simple on Lyftup. You upload your media after your media has been uploaded. The first thing the system is going to do is it's going to actually detect and then track all of the faces that it's found in the media. And then it's going to prompt the user to go ahead and actually label the faces that it's found. And once labeled, LipDub then understands identity. That process of uploading media and tagging can take, let's say for an interview that's an hour, probably takes about 20 minutes of pre-processing time. And then once you have a processed video, all you're doing is there's a training step in between.

13:19So what our system does is we actually train on the speakers. And the reason we train is purely for texture. To get that heightened texture. That's the longest part of our process computationally. It used to take 10 hours. Now we're down to two and we're continuing to try and drive that time down. And then once you have labeled speakers that are trained up, all you're doing is you are associating new audio files to those speakers. So it's a simple drag and drop. So you'll take, you'll take as an example, you'll go into deep dub or 11 labs or whatever, one of these audio tools you'll go and you'll create the Mandarin voice track for Craig.

14:06And then you will simply associate that to Craig and click go. And within 10 minutes or so, you'll have, you'll have a video of you speaking Mandarin and the same would go for me. So that's the general flow on the platform. What the platform is doing is manipulating the pixels around the mouth, presumably. I would guess that's frame by frame. Is it done with a patch? How large an area of pixels are replaced? and how is it associated with the lips closing or opening with the audio? Yeah, so what we're generating, it's pretty much everything from below the eyes is a reconstruction that's based on the audio.

15:00We tinkered over time and that continues to evolve what we mask out and what we generate. But today in the state of the product, pretty much everything from below the eyes is a reconstruction. In terms of how the system works, there's the obvious aspects of it that most people could figure out. Again, the relationship between phonemes and bisemes, phonemes being the sound, the oohs, the ahs, right? There's only a finite number of phonemes. There's then bisemes that correlate to those phonemes and there's that mapping being done. but then that's where it really starts to go into like tail end problems because it's not just about lip shapes right like one one of the realizations that we had pretty early on was the importance of mouth internals so much of how we speak actually happens it's not our lips it's our tongue it's our teeth some words are produced almost entirely with our tongues so you could have two very similar our mouth shapes, but different tongue and teeth positioning generate different sounds.

16:05That was a huge challenge for us to figure out, like, how do we do mouth internals properly? And then it just keeps going. And then it's all about how do you personalize? How do you make sure that what I'm reconstructing doesn't just look like any set of lips or a random set of lips or a proxy of the lips? How do you make it look exactly like the speaker? And then you just keep going down this long tail end of problems. Yeah. And we talked about the consumer grade products that are out there. And we're going to do this in Chinese. I have an audience in China. Yeah. Do the Chinese have a similar solution?

16:48Because oftentimes they're competing at the cutting edge with the U.S. solutions. Yeah, not really. Look, there's quite a few products now and it's validating for us because we feel like in many ways we started this category. There was one company that predated us on lip sync, but they didn't focus on automation and that was so important to us. Not automation for automation sake. We always felt that even if we could lip sync, if it took too long or if it costs too much, it would just limit the accessibility for most use cases. So in terms of the first folks globally to actually automate something that operated at this quality level, we really felt like we launched this category.

17:45And yeah, certainly now we see a bunch of folks coming in and call them like fast follower companies. The difference is most of these companies are just wrappers. Yeah. They're just wrapping around open source. And so of course they're inherently limited by open source and what open source gets you. We started with open source two years ago and just realized that it's not even getting us close to where we need to be. But no, we haven't seen any Chinese competitors. The reason I ask about China, there was a famous video that I think it was SenseTime or iFlytech, I don't remember which Chinese company, put out during Trump's visit to China in which he spoke in Chinese, which kind of blew everybody away at the time.

18:37That was a voice cloning, but the lip sync wasn't there. and so I was wondering whether the Chinese have addressed the lip-syncing portion. How long does the process take? Is there a metric like so many minutes or hours for every minute or hour of video that you're lip-syncing or is it variable depending on how, as you said, dynamic the scene is? Yeah. Roughly speaking, every new minute of content that you want to generate on platform, it can take anywhere from right now 10 to 20 minutes. Although it's not linear. it's not as though if you run an hour of content through the system that it gets faster as it moves through the content but because we've built everything in a scalable manner so all all these processes can take place in parallel so as an example if we were to lift up this interview into 10 languages, you'd be able to generate all 10 new videos simultaneously on the cloud.

20:05And you could probably guesstimate, yeah, it would probably average out over the course of an hour to probably around 10 minutes per minute. Doesn't include the training. The training is that kind of, you have to do it. It's two hours. You do it once. You don't have to do that per language. You just do that once just to really understand textures. And then, yeah, you're looking at probably around 10 minutes per minute. Yeah. And the cost, is there, how do you price this? It's a self-serve platform, right? You walk the clients through it, but once they understand the platform, they should be able to do it themselves.

20:49is do you is it a subscription model do you charge by minute or how do you do that yeah you're exactly right so it's a subscription model the way it works is you purchase credits up front on the platform you can either purchase credits monthly or you can purchase credits annually if you purchase your credits monthly it's a kind of a use it or lose it model where you'll get that allocation of credits for the month and whatever goes unused will expire at the end of the month. If you pay for your credits annually, you'll get all of your annual credits upfront and you'll have the flexibility to use those credits whenever you need during the year.

21:31And the price of the credit is$1. The difference is the number of credits you consume will depend on the activity that you are running on the platforms. Generating a 1080p output video will consume less credits than generating a 4K video, as an example. Yeah. And who's the primary use case? You guys built this for Hollywood, but as voice cloning and real-time translation, develops. It seems to me that this solution will be in higher and higher demand across all kinds of domains. Yeah, that's why we're excited for LipDub to be valuable, right? There needs to be some new dubbed audio that you're trying to associate.

22:34And historically, dubbing's been this very manual, very expensive process that's really only been leveraged by Hollywood and advertisers for the most part. And now we ask ourselves, okay, now that dubbing's becoming this very affordable, very accessible task, how much of the world's content is about to get dubbed. Right now, only 1 % of the world's video content has been dubbed. But again, that's predicated on the idea that dubbing has always been this very manual, very expensive task. If you can now dub for pennies per minute, what percentage of the world's internet content is about to get dubbed?

23:28And we feel very strongly that anything that's dubbed deserves to be lip synced so in terms of circling back to your original question who's the user that's a constant conversation really excited by the youtube market like really truly excited there's a number of proof points out there at this point around from those kind of early adopter innovative youtube channels that went ahead and dubbed like Mr. Beast. The statistics that have come back from that, this kind of two-year experiment that he's run have shown that there's a huge global appetite for his content. Yeah, actually, I didn't realize that Mr.

24:18Beast had dubbed his videos. What language has he dubbed into? He started with 15 and he's going to 30. Wow. Yeah. Yeah. And he started his experiment before the AI audio tech existed. So originally he was paying traditional dubbing studios to do the work as a way to validate the experiment. Does this work? If I dub my channels, is my content going to perform well? And he's, it's not as though he posts all of his performance metrics, but he's posted certain months as an example, where he gives you a breakdown of where his views are coming from and 50 plus percent of his views uh come internationally in some of these months from dubbed content or just internationally watching it in english your business is ready to launch but what's the most important thing to do before those doors open getting more social media followers, or actually legitimizing and protecting the business you've been busy building.

25:27Make it official with LegalZoom. As a business owner, traditional legal services come with major sticker shock. Getting registered or talking to an attorney shouldn't have to cost you so much. Well, thankfully, today's sponsor, LegalZoom, created a better way to start and stay in business from initial formation to one-to-one legal consultation. I've used LegalZoom for over 10 years, and it's kept me compliant. They have everything you need to launch, run, and protect your business all in one place. Setting up your business properly and remaining compliant are things you want to get right from the get-go, but you don't have to strain your brain or wallet.

26:15LegalZoom saves you from wasting hours making sense of the legal stuff. At LegalZoom.com, you can take care of business legal needs in just a few clicks. And if you need some hands-on help, their network of experienced attorneys from around the country has your back. Launch, run, and protect your business to make it official today at LegalZoom.com and use promo code SMITH10 to get 10 % off any LegalZoom business formation product, excluding subscriptions and renewals. This offer expires at the end of the year. Get everything you need from setup to success at LegalZoom.com and use promo code Smith10.

27:06That's S-M-I-T-H-1-0. LegalZoom.com and use promo code Smith10. S-M-I-T-H-1-0. Duh. Wow. Are you working with him? Or is this something that YouTube could integrate into YouTube Studios so that people at a click of a button could lip sync dubbed audio. Yeah, so we are working with Mr. Beast. We're starting to get into the exploration of working on some lip sync. And there's a number of other major YouTubers that we've brought onto the platform recently that are either A, already have dubbed content or are just seeing the trends and want to jump on and start to localize their channels because it really is.

28:01It does represent relatively speaking, low hanging fruit for them. You've already invested in the content. And now it's just about how do I, how do I squeeze the lemon? How do I extend or how do I optimize the value of the content that I've already created? Localizing is, it's a great way to do that. it's not the only market but i'm just particularly excited about that market because i really do believe like the world like there's no reason that we only watch influencers that speak our language people are making interesting content everywhere and i think you can just look across the whole media spectrum and there is this demand i'm currently as an example i'm really into show gun.

28:50I think it's fantastic. Squid Game was fantastic. We're also working with some YouTubers now that are major influencers in other parts of the world that are really interested in tapping into North America for the first time. And captions just, which is the way they've historically done it, is just, it's not very engaging, right? It's a second tier way to view that person's content. And now all of a sudden you're able to have a YouTube channel where you can give. Every country in the world, this quote unquote first class viewing experience as if it was made for you. And I'm really excited by that.

29:34The other markets that are leaning in right now are advertising is a really big one, both digital marketing and TV broadcast. And a lot of the clients signing up are either the advertising agencies or their video production companies. We just did a TV commercial with probably my favorite tech brand ever. And that's going to come out shortly. We'll be able to talk about that shortly, but that was really exciting. And then I mentioned it earlier, online education, whether it be for your employees, let's say you're a multinational that has employees across the world being able to communicate with your international workforce or folks that are selling courses that want to go and enter new markets, right?

30:25We have a few of those right now, folks that have meaningful course loads, very successful companies, but only successful in their region. And now they're looking at lipdub as a mechanism to enter new markets and grow their businesses. So it's definitely a lot to keep track of, but I think in a good way, I think in an exciting way. So it goes both ways. There are people producing content in English that want to reach non-English speaking markets, But there's a tremendous amount of content. I spent much of my life in China in Chinese that the English-speaking world never sees. And that's frankly one of the reasons I believe there's such a gap in understanding between the two countries, because people just aren't exposed to Chinese life.

31:27And so is there a dominant language? Is most of the content, most of the market, from your view, translating English language content into other languages? And which other language do you see as being the primary target or what top three? Or do you see the market really translating foreign language content into English to reach the English speaking world? It really is both. I don't see a dominant trend either way. For Hollywood specifically, their initial use case and what they're particularly most interested in is foreign to English. Likely just because we as English speakers, we've had, we have no patience.

32:17That's other markets like Germany or France that have grown up on dubs. So the idea of the lips not being synced is it's not like it's ideal or optimal, but at least they've grown up with it. Whereas we have very little patients and we're very attuned when there's that issue. So Hollywood's certainly interested in the foreign to English. But when it comes to advertising, online education, YouTube, we really see it all. see all the major European languages, German, French, Italian, see a lot of Indian languages, like Hindi, Mandarin is a huge one. So we're really seeing, we haven't seen any one particular trend stick out.

33:07It just feels like everyone wants to better communicate with everyone. And how real-time could this get? Is it conceivable that eventually you'll be able to sync and dub live streaming content with some delay? Yeah, certainly conceivable. The challenge typically when it comes to working with real time is there's typically some quality trade-off that you're making. But a lot of times now with technology, technology makes old trade-offs disappear. So we certainly are interested in that as a, as a future for like future development. Because obviously if you can do real time, you open up a lot of interesting use cases.

34:08And at that point it really becomes an important cog in the universal translation machine. Just the idea that I could speak to a coworker in China and have an experience where I can connect with that person in a way that I just never historically could is obviously really interesting. And then you just have a lot of content that is inherently live content. A lot of broadcast is live. Although we do see some broadcast use cases going on the platform. As an example, there's a few companies now doing cricket analysis for all the different state languages in India. Those are, yeah, which I think is really cool.

34:56India is a great market for this. India is like one of the best markets because there's so many dialects spoken and typically you either have to create content for each dialect or some dialects just don't get a great viewing experience of content. So very bullish on India for this tack. One of the challenges isn't simply the lips opening and closing or the position of the teeth or tongue, but also the phrasing, because something in translation can take longer to say than in English or vice versa. How do you deal with that? Yeah, it's a great point. These are the two, I'd say limiting factors in most of the AI audio software.

Read the full transcript

35:50So the first is translation accuracy. Some languages translate at a much higher accuracy rate than others, which is something I think it's, it certainly needs to be solved. And the other, which is an even harder solve is colloquialisms and slang. And I, I, I'm confident that that's a subset of the translation accuracy issue, but both of those things are real problems. And that's why in most of these AI audio systems, right, you are given the ability to go in and edit the retargeted script, but that requires then somebody that speaks the language to go in and do that work. So that just makes the systems more difficult to get value out of, right?

36:47If in order for me to perfectly translate a video into 10 languages, if I need a speaker from each of those target languages to vet the translations coming out of those audio platforms it's not to say it's not doable it's just somewhat annoying and logistically challenging so that's certainly an issue that exists today in the platforms and the people that are using ai audio most of them are going to those lengths to actually do that work and get people to those languages the other issue that you mentioned is timing which is another it's a limiting factor in the systems, there's kind of two things those programs do.

37:31They will shrink and they will shrink and stretch audio to match. But that approach can lead to interesting results at times, right? If you're listening to an audio where part of that audio feels sped up and then it slows down, there's a fine line between an acceptable viewing experience and something that just ends up totally distracting you. Your work around there though goes back to the script editing piece. If you have English content that's going into Spanish and that Spanish audio out of the box is 15 seconds but the English was 10 seconds, yeah you can rely on that automated slow down speed up or you can actually go in and tweak the spanish script take out some words tweak it a little bit exactly what hollywood does by the way yeah again that's cumbersome but it but it sounds like it's it sounds like that could be automated in the translation and voice generation if you put in the timestamps that you're trying to match.

38:53It seems there are many ways to translate a sentence generally. And it seems to me you could automate that so that it's translated to come close to the time used to say the same phrase in English or whatever language you're dubbing from. Yeah, I think that could be an interesting way to do it. I think what you're proposing is potentially it's, Hey, here's a few iterations that capture the sentiment of what that original script said. And this one's shorter and fits your video better. And this one's kind of the verbatim, but it's too long. And I, I'm not suggesting that any of these are unsolvable problems.

39:35Yeah. These are just the limitations today where some people walk into the platforms and expect perfection. both like lip syncing software like what we have it's pretty magical the ai audio software is pretty magical but magical doesn't mean perfect and magical doesn't mean that there's no work entailed the obvious question is the potential for misuse so how do you guys think about that the better this technology gets the harder it's going to be to distinguish whether something is dubbed or and lip synced you'll easily be able to have people saying things that they didn't say and how do you think about that and are are there any controls or anything built into the platform or that you're considering building into the platform to to police that kind of misuse it's something we talk about quite often.

40:45We built this to help ultimately to help the world better communicate, not to topple democracy. So there's a few things that we do. One thing we do is we ensure that whoever you're Lipt-dubbing, you have to click in the platform that you actually have permission to Lipt-dub that individual. We also spot check everything that runs through the platform. If we see misuse, if we see a celebrity that's promoting something that we know they didn't promote, you're banned from the platform for life. Those efforts are time consuming, but we think necessary. I think a lot of this will eventually come down to great collaboration between those facilitating AI-generated content and the distribution platforms for that content, be it social, YouTube, Facebook, you name it.

41:48There are ways to tag these pieces of content with metadata. There's already, I forget which platform that's already coming out and making sure that any piece of AI generated video content is going to be labeled as such. And I think that's crucial because to your point, it's, we're already, we're already past the uncanny valley. If we don't, if everybody doesn't start to work together on this, then I think the negatives of all this AI, gen AI tech are going to be real, like really detrimental to society. And I don't think anybody wants that. I don't think that's why anybody built these tools in the first place.

42:34Like they built these tools to foster creativity, to give people the ability, new capabilities to do things that they just never fathom being able to do, power the individual. but yeah without close collaboration i think things could go haywire there's been a lot of research into digital watermarking or you know embedding some patterns in the pixels that aren't visible to the human eye that could be read by by machine are you guys talking to any researchers about those kinds of solutions. That's exactly what I'm referring to. Yeah. This digital watermarking. I'm not the person on our team that actually drives those discussions, as I'm not an engineer and I'm not technical enough.

43:29But certainly that's what we aspire to deliver on, is something that can't be manipulated by humans, somewhat defeatist if it can be. We want these digital watermarks to be permanent because again, we're here to help people communicate. We're not here to foster misinformation. So yeah, it's super important, but it's like anything, right? Like any new technology, most new innovation, there's the good and the bad that comes with it. Are there use cases out there that people can look at? any Hollywood examples that have used your tech or YouTube examples? I guess you were saying Mr. Beast, but that people could take a look at.

44:18Yeah, in Hollywood right now, honestly, the majority of the work that we're doing has been what they call ADR. So ADR is where you want to change something in your film. Something's been shot where you're like, I didn't love how the actor read that line of dialogue, or I'd love to change that line of dialogue, which predating our software would typically necessitate a very expensive reshoot. So that's the kind of work we're doing with Hollywood today. Candidly, the product needs a cost structure that works for Hollywood to do this work. A lot of our YouTubers are just getting started right now.

44:57So you'll see them on platform very shortly, and you'll start to see their content popping up. And then there's advertisements, right? We just did a great David Beckham spot for Lays, where we changed, where we localized, where we localized into different languages. I just mentioned that major tech brand that we just did, basically three campaigns, four into eight languages. That'll be dropping really soon. And some of our clients are, again, doing work on behalf of brands, either for their digital marketing or e-learning efforts, where the metrics are pretty amazing in terms of like viewership and engagement rates.

45:44But those aren't metrics that we can necessarily share. How did you put together the team to do this? Who were the founders? just what was the origin story you talked previously about why you guys started and what you started with but i'm always amazed you're a lawyer that you were able to pull together these kinds of engineers are very expensive and in high demand uh yeah just talk about that a little bit. I'm surprised too. So it's not just you. Yeah. It all honestly started like this was my previous software company. We didn't do AI work. And so AI work is, it's a unique, it's a unique animal.

46:44And one thing that I learned very quickly in setting up like the first iteration of this team was it's certainly not a game of like where quality is trumped by quantity there's many people that are researchers you can hire a room full of average researchers and you will get a hundred reasons back as to why a problem is not solvable whereas you can hire one incredible researcher and they will give you the answer to your problem. And after the first iteration of Mars AI, I recognized very quickly that if this was going to be a serious program and if we were going to develop world-class product, we needed world-class research.

47:36And it was really hard to get Daniel Cohen-Or. So Daniel is the world's leading SIGGRAPH. He's the number one most SIGGRAPH published contributor in the world. His lab out of Tel Aviv University is world renowned. They pushed the pace of it. They started as a graphics lab 30 years ago, but 10 years ago, were one of the first groups globally that started to ask themselves the question of like, how is deep learning going to implicate graphics? and since then the papers that are released by their lab and the students that have come out of that lab are some of the most important people in this world of graphics today and so I really wanted Danny because he just fit very perfectly.

48:26I ended up standing an advisory board some really great Canadian professors across different institutions. So that was my step one, was building this advisory board and then basically going through the networks of the advisory board.

48:46And through those connections, I found my way to a number of great candidates. And some of those interviewing processes were super lengthy. And eventually I really had my eyes set on Danny and it took eight months to sign him. But I think he was excited by the vision because it's so aligned with what his lab works on. And then once Danny signed on as chief scientist, we locked into his longtime collaborator, who's an assistant professor at SFU here in Canada, which is a top computing school in North America for graphics. And his name is Ali Madhavi Amiri. So we got our chief scientist. We got our director of research.

49:43research and that ended up being like, that was the bedrock. Once we had those people, you had an environment where the idea of democratizing VFX is in and of itself pretty attractive. And on top of that, you're able to go and speak to researchers and basically tell them, look, not only do you get to work on these really interesting problems, but you get to do so with some of the most talented people doing it globally in the space. And then it starts to build on itself. And, but yeah, it was, it all started with that advisory board. And have you raised capital before pulling together the advisory board, or did you pull together the advisory board first and then raise?

50:37What did we do first? I want to say that we pulled the advisory board together prior to raising the money. Yeah, that makes sense. We were talking at the beginning about getting the camera's position so that you're not looking away while you're speaking because you're looking at your interlocutor on the screen. Yeah. Have you guys talked about manipulating the iris position? I know that exists, but it's... Nvidia put out something along those lines. Yeah. Look, we're, we're certainly not, we're not done when it comes to product development. We definitely think that Lyftab is a great product. We think it has a place in the world.

51:24Uh, for us it's a launch pad. I think we're going to be doing a lot of a new product development as the company grows all along the lines of how do you empower the individual from a creative standpoint by giving them pretty much like access to different VFX applications that used to take teams of artists. And our whole thing is have it be something where the individual can come in and just tap into that kind of creativity. So like things along those lines, but we're certainly not done. I don't know if Iris change will be the thing just because like you said, it's been done. But we've got some, we've got some cool ideas that we're floating around, but we really are trying to remain focused at least for now on like seeing through this lift up mission, which is far from over, but it won't be the last, it won't be the last product that we put out.

52:21That's for sure.

From the publisher

Launch, run, and protect your business to make it official TODAY at https://www.legalzoom.com and use promo code EYEONAI to get 10% off any LegalZoom business formation product excluding subscriptions and renewals.

 

 

In this episode of the Eye on AI podcast, Matt Panousis, Co-Founder & COO @ LipDub AI, shares the journey behind the platforms a cutting-edge AI-powered lip-syncing tool revolutionizing the dubbing industry.

 

Matt joins Craig to explore the future of content localization, the evolution of LipDub, and how generative AI is transforming how audiences experience video content worldwide.

 

As a leader in AI and visual effects, Matt discusses how LipDub integrates advanced AI to enable real-time lip-syncing for dubbed audio, making Hollywood-quality dubbing affordable and accessible. From processing hours of video footage to handling complex, dynamic scenes, LipDub is empowering content creators to reach global audiences like never before.

 

We explore the challenges of creating realistic lip movements, the importance of mouth internals, textures, and fidelity, and how tools like LipDub stand apart from competitors like Rasque AI. Matt also introduces Vanity AI, LipDub’s sister technology for automated digital makeup and de-aging, used by MARZ for high-end visual effects work.

 

Learn how LipDub is redefining the possibilities of content localization with AI and what the future holds as this technology evolves.

 

Don’t miss this deep dive into AI-powered innovation with Matt Panousis. Like, subscribe, and hit the notification bell for more episodes exploring cutting-edge advancements in AI!

 

Stay Updated:

Craig Smith Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI

 

 

(00:00) Introduction & Matt Panousis Background
(02:45) Origins of MARZ and LipDub
(04:20) LipDub vs. Competitors
(06:15) Integration with Dubbing Solutions
(08:00) LipDub Technical Process
(10:20) Competition
(12:20) Time and Cost
(14:00) Target Markets and Use Cases
(19:00) Future Developments
(22:00) Translation and Timing Challenges
(28:00) Ethical Concerns and Misuse
(30:30) Digital Watermarking
(32:25) Real-World Examples and Future Products
(34:00) Team Building and Research
(37:00) Building the Advisory Board
(38:10) Securing Research Talent
(40:00) Funding and the Board
(40:35) Future Products

More from Eye On A.I.

All 266 episodes
#224 Matt Panousis: How AI Is Disrupting The Dubbing Industry (LipDub AI)Eye On A.I. · 52 min
Listen in VO