In short
NVIDIA AI Podcast - Episode 202: Deepdub’s Ofir Krakowski on Redefining Dubbing from Hollywood to Bollywood
Podcast Overview
- Title: NVIDIA AI Podcast
- Description: The podcast explores how innovative technologies are transforming various sectors, including entertainment, through interviews with industry leaders and pioneers.
- Episode Title: Deepdub’s Ofir Krakowski on Redefining Dubbing from Hollywood to Bollywood
- Release Date: August 30, 2023
- Link: [NVIDIA AI Podcast](https://ai-podcast.nvidia.com/)
Episode Summary In this episode, host Noah Kravitz interviews Ofir Krakowski, co-founder and CEO of Deepdub, an Israeli startup focused on revolutionizing the dubbing and translation process in film and television through generative AI.
Key Themes
- Global Entertainment: The episode discusses the broadening scope of TV and film production beyond traditional hubs like Hollywood and Bollywood, emphasizing the growing importance of dubbing and localization.
- Challenges in Traditional Dubbing: Traditional dubbing methods are often slow, expensive, and fail to capture nuances in language, such as idioms and cultural references.
- Deepdub’s Solution: Deepdub utilizes AI to streamline the dubbing process, enabling efficient translation, voice generation, and audio mixing while ensuring human oversight to maintain quality.
Key Concepts and Discussions
The Problem with Traditional Dubbing
- Inefficiency: Traditional dubbing can take months and involves numerous actors for diverse characters.
- Cultural Nuances: Current technologies often overlook subtleties in language, leaving essential elements like jokes and idioms lost.
Deepdub's Innovative Approach
- AI-Driven Platform:
- A web-based platform that automates translation, voice generation, and audio mixing.
- Allows interaction with sophisticated AI models to enhance efficiency.
- Human Oversight:
- Despite AI capabilities, human verification is essential for ensuring natural sound and emotional accuracy.
- Lip Sync Technology: Deepdub is developing technology to match lip movements to dubbed voices, enhancing the visual authenticity of dubbed content.
Future Aspirations
- Global Accessibility: Krakowski envisions a future where language barriers are minimized, allowing global audiences to enjoy diverse storytelling.
- Expanding Use Cases:
- Potential applications in diverse fields such as advertising and e-learning.
- Ability to create localized content for various cultural contexts.
Challenges Ahead
- Translation Nuances: Achieving perfect translation for complex phrases and jokes remains a significant challenge.
- Real-Time Translation: Future goals include enabling real-time dubbing for live content, such as news and sports.
Conclusion Ofir Krakowski emphasizes the transformative potential of Deepdub’s technology in democratizing access to entertainment and education globally. By breaking down language barriers, Deepdub aims to enrich cross-cultural understanding and storytelling.
Additional Resources
- Deepdub Website: [Deepdub.ai](https://deepdub.ai)
- Follow on Social Media: Check the official website for links to social media platforms and newsletters for updates.
---
This episode highlights the intersection of AI and entertainment, focusing on how innovations like those from Deepdub are set to redefine how global audiences consume content.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:10Hello, and welcome to the NVIDIA AI Podcast. I'm your host, Noah Kravitz. One of the things that modern AI is very good at is translation. Large language models showcase AI's prowess at translating written text between languages. But machine translation has been a useful real-world tool for some time now. That said, translation apps and auto-generated video captions are great. But what if we could use AI to automate overdubbing of voice content into different languages? High-quality, low-cost dubbing could do wonders to open up more of the world's audio and video content to broader audiences across language barriers.
0:45Our guest today has been working on just this problem at the Israeli startup he founded with his brother. Ophir Krakowski is the co-founder and CEO of DeepDub, an AI-driven dubbing solution used in Hollywood movies that are currently showing in theaters globally. DeepDub aims to bridge the language barrier and cultural gap of entertainment experiences through high-quality localization at scale. The company's deep learning-powered dubbing solution helps studios, broadcasters, and distributors with all of their localization needs. From translation and adaption, continuing to dialogue creation and finishing with the final mix.
1:18Ophir is here to tell us more about DeepDub and how generative AI can help connect the world by breaking down language barriers. So let's get right to it. Ophir Krakowski, welcome, and thank you so much for joining the NVIDIA AI podcast. Thank you, Noah. Thank you for inviting me, and I'm really looking forward to talking to you. Likewise, it's our pleasure to have you. So let's get started by hearing a little bit about the story behind DeepDub. If you would, take a moment, tell the audience about the company and how it got started. As some may know, I've been the head of the machinery and innovation department in the Israeli Air Force.
1:54So I've been used to doing stuff that are impacting the life of people in Israel. But as I got out of the Israeli Air Force, I thought that I can do something with my knowledge to impact humanity in essence. And I searched what I'm going to do. And I founded this company with my younger brother, Nir. And in essence, what we found is that content is always created in one language. And when you want to transfer it or to reach more audiences, you are facing a language barrier and a cultural barrier. And the current language models currently don't support some of the use cases like jokes, idioms and places and even regular phrasing that you use on jargons that go on the streets nowadays.
2:54And in essence, they don't support the delicate intricacies of a language. And in this essence, we thought we can develop something which was very ambitious three or four years ago to take this problem and develop a technology and a product that would enable everybody to use it. We figure out that if we can do a theatrical for one of the Hollywood studios, then we can serve everybody down the line. Right. Listening to you talk about it makes me think of the old, it makes me think about how old I am, but the old sort of trope, at least in America, of kung fu movies that were overdubbed. And you would see on screen the actor's lips would start moving and then maybe a second later the dub would come in and it wouldn't match up at all.
3:45And to your point, there's so much more than just the words and the translation, which there are plenty of nuances in language, but you get into, as you said, idioms and slang and just the cultural backdrop of how the words are being spoken. And there's a lot to it that, from my standpoint anyway, I believe you when you say it's tricky. So how did you get started with it? Now, you mentioned both you and your brother come from backgrounds working with machine learning and AI. In what sounds like quite a different context. How did you make the move from you were serving in the Air Force, your brother was working at an intelligence agency in Israel?
4:26Yes. So how did you get the idea to move from that? And then how did the actual transition take place to starting DeepDep? So both of us are a true lover of the magic of creation of content. So this is a part of like an hobby. Yeah, yeah. And building this company is just combining our abilities with and impacted in a good way that content can be accessed to large audiences. And so combining the love to creation of content with the love to technology, this is something that led us to found this company. And were either of you working specifically on synthesized voice and language translation problems?
5:10Or was it kind of a move, using your background in AI and deep learning, but a move into kind of new territory for you? So this is a very interesting question you ask, because most of the technology of creating voice was dominated by four big companies three or four years ago by the big, you know, meta, Amazon, Google and Microsoft. So in an essence, the knowledge was not out there. We had totally to invent everything from ground up and take the most recent research and incorporate in the company people that are from these companies and to build the know-how. We actually not only build the product, but we also built the AI infrastructures to train new models.
6:06So in essence, this is what we did. And we built something that, as I told you, we built something that can screenshot and allow technology to create a voice, a human natural voice, which the right pronunciation with the little intricacies of the human voice. So in essence, you cannot differentiate between a human voice and a machine-generated voice. Right, which brings up some questions that I'll put a pin in to ask you a little bit later. But I want to ask you now about how does the solution work when you work with, and I don't know what the best way to get into this is, so feel free to take a different tack.
6:50But when you're working with a client, say a Hollywood studio that wants to leverage your tech to dub a piece of content, a movie into different languages, what's that process like? How does the solution work? Let's talk about the numbers. Let's take, for example, a TV series. So we'll have like 10 episodes. It will take you like three months to dub it using humans because in a full season, you'll have like 100 characters in an average, like a TV, a drama TV series. You'll have 100 characters. You have to bring into the studio a lot of people. Right. You have to make sure of diversity, of the AI.
7:32So you need to take care of a lot of things when you are trying to dub the content or to make it available to other audiences. And then you need to do it across 32 languages. So this is not a technology problem. It's a huge project management problem, right? So we thought that we can make this process very efficient by introducing efficiencies in the entire process, not just creating the voices, but also in the translation, the adaptation and the mix itself. So just to make the audience aware, the process is very convoluted because it takes a lot of phases. In the first phase, you just get from the studio, as you just mentioned, if the customers bring us a movie or a TV series, then we get the video and the audio.
8:26And the audio is, most of the cases, it would be in English. And then you have to transcribe it because you need to know who said what. And then you have to translate it. And then you need to adapt it because, for example, a joke that it will work in a different language. It's not a straight translation like the regular translation tools that you have. And then you have to create voices. And after you create the voices, you need to mix the audio with the music and effects and glue it back together to the video. And this you need to do in every language. Right. Like in 32 languages. 32, yeah. Yeah.
9:04So taking this into account, it's a very convoluted process. A lot of time, a lot of people involved, a lot of different parts of the process. Absolutely. Right. So actually what we've developed is a platform, a web-based platform that enables people to interact with a very sophisticated AI models that enables them to do each part of the process in a very fast manner. And this is for the high end. Maybe later I will elaborate more how this can affect more customers because we are currently developing something that will enable other customers to work on the platform and not just studios. No just studios.
9:48But when you work with a student, this we need to understand a studio doesn't want a machine learning tool. It wants a white glove service because it wants it to be in the highest quality. And the best AI model currently available cannot get past 95 % of accuracy. Even the best chat GPT have mistakes. So you need the human in the loop. And this is why we have developed the platform. So it's a kind of a Adobe Premiere kind of a tool, but for creating localization and creation of voices, which is very simple. So it will take someone one to two hours of very simple training to deliver a drama episode.
10:36Would your platform take care of every part of the process? You mentioned the transcribing all the way through to Final Mix. Yes, definitely. You know, we understood at the first stages that nobody cares of creating voices. Everybody cares on the end product. You want the end product. You want the localized video. You know, just bringing you the voices. You need to do all the other work that I just talked about. Because even if I created the content, for example, if I have created a podcast, I know English, I know Hebrew, but I don't know Chinese. Right. So I don't know how the translation went.
11:20You know, is it good? Is it how it works? So I need someone to help me curate. And so let's get into the human in the loop. What is the human doing when they're localizing an episode of the TV drama? If I want to localize that using DeepDub, what are the steps that the platform takes care of and where does the human come in? I have a million questions, so I'll let you talk through it and I'll jump in. Okay, so in essence, the first cut is done automatically by the platform. Okay. But then a human comes in and verifies that everything works well. So it verifies the translation. It verifies the generation of the voices.
12:04Like, for example, do the emotions are the emotions that needs to be at the target language? You know, in different languages, question is asked in different ways. So in an essence, you need to understand this and a model sometimes know how to do it and sometimes mistake on it. You know, since we're recording everything, so we're recording the curation on the text, curation of the voices. We are making the machine learn from these mistakes and become better and better. So over time, the intervention of human in the process will be less and less needed. What's the most difficult part of the process for the machines to take care of?
12:49Is it, you know, is it the getting nuances of the language? Is it getting the voices and the emotion right? And I guess I'm asking both in terms of what was tricky for you and your team or is ongoing and building and refining the platform. And just from sort of a purely technical aspect, is there one part of the process that's just a harder problem to solve objectively? You know, I believe that building a machine that can support a wide range of emotions from text, this was to us the first problem that we tackled. And it took us about two years to develop something that will support a wide range of emotions that can support a theatrical.
13:36Right. Which is not like a podcast or audible, which is mostly a narrow range of emotions. Yeah. So in a theatrical, you have somebody screaming like from the heart or talking while eating. This is a different, it's funny, but it's a different voice. Like you don't understand it, but the machine looks at it as a different voice. It's like digitally is a different voice, but we support it now. So I cannot tell you this is not a difficult question, a difficult issue. But the most difficult issue is I think that is currently not solved is the translation part. Translation part, I think, you know, as I believe it, as I know the material, I think that getting a 100 % translation from the machine that will 100%, for example, a joke, it will take some time.
14:28Yeah. Does your system analyze the video content as well to pick up on facial expressions or body languages or even patterns in the way characters are situated in an episode might give some cues as to what's happening, what the subtext is? Or is it strictly looking at the audio? So we're looking at all aspects of, it's called multimodality. It's not looking on the video, the audio, and the text. The fact that we're using LLMs is that the actual LLM understands the emotions and understands the text. It's kind of a, I don't know, this is a very simple way to explain it. But it's actually understanding the text.
15:15And it can convey the emotion to different languages. the emotion that is in the text. In essence, we even use the video. We have the technology currently where it's not a product that is out there, but we have it in our labs that changes the lips just in the places where you cannot do it with words. If you are translating from German into English and you end up translating the word no, so in English, no is an open mouth, but in German it's nay, it's a closed mouth. And it's a short word and you don't have anything to do. So you have to struggle when translating it. So in an essence, in those places, we call it the last mile solution.
16:02You would use a change of the video, but this would be available soon. This kind of solution also on our platform. We currently, it's currently supporting 1K. And as we are aiming always to deliver first to theatrical, which should support 8K or 4K. So as we support this, we'll have this enabled also to other customers. So currently, it's audio dubbing that you're doing, but you have been thinking about it. And you just alluded to the solution that you're working on that does actually, I don't know if the right word is generate or recreate some of the video content as well to match a new audio.
16:46Right. Yeah. Right. It's currently the General AVI model support all of this. And this is how they are going to impact the entertainment industry because it's not just the audio. But I think the audio part is one of the, you know, audio translation or voiceover. It's a tradition of 100 years. Yes. like the from the Mussolini times it was started starting to dub content right until this day they do it but most of the content around the world is we need to understand it's not dubbed so it's not accessible to most of the the people around the world you know I don't know if this podcast is localized in that so far as I know no I mean we show up in in some of the international you know the analytics right we have listeners internationally but uh so far as I know it's not localized so um right so most of most of probably most of the the you know latin american audiences which some of them don't know english very well this podcast is actually not accessible to them they just tune in because they like the way my haircut looks on the radio so that's you know that's that's why they're tuning in and your voice i guess i i haven't seen obviously i haven't seen the solution that you mentioned that's not out yet, but I have seen some early attempts from other companies at doing what you were describing, changing.
18:11I'm pointing, nobody can see me on listening to the show, but I'm pointing to my mouth, changing the way that lips look to accommodate different audio overdubs. And they look really creepy to me. The ones that I've seen, and probably it's new technology, it's not there yet, but it's just this weird, uncanny thing of the rest of your face, you know, not moving or just having a different expression, but then the lips are, you know, clearly doing something different and it matches the words, but doesn't match the rest of the person's face. I can only imagine how difficult and tricky that technology is to develop and get right.
18:52But it's an interesting time for millions here to say the least yes yes but but but you know with audio itself you can solve a lot of problem and since most of the audience are currently used to get the dubs in a way that it doesn't interfere to them if there is a little except for the u.s audience but other audiences around the world which are used to to dubbing or voiceover for a lot of years they're using to consume it, you know, in the way that it is right now, only the audio. But for English audiences, this is very important because, you know, audiences that are not used to dubbing content.
19:30And this is very interesting because, you know, since Netflix started to dub content, international content in English, there are more and more Americans that are exposed to international content, which I think is very good because they are now open to more cultures around the world. In fact, it's interesting because most of the content we've dubbed is into English. So we have hundreds of hours of international content dubbed into English. Oh, interesting. Okay. Yeah. My guest today is Ophir Krakowski. Ophir is co-founder with his brother and CEO of DeepDub.ai, an AI-driven dubbing solution
20:18was talking about starting with theatrical releases, working with big studios, and then kind of trickling down as the technology becomes more mature and I would imagine less expensive to use going forward. So we've been talking about the company, about the technology and the importance, and as you were just saying, unlocking not just all of the English language content exported from the US to other cultures, but also kind of moving back the other way with streaming platforms. And as you mentioned, there are a couple of Japanese shows on Netflix that my family and I watch from time to time. And we just watch them with, actually, now that I'm thinking about it, some are subtitled, and then some do have English, at least the narration is in English.
21:03And so to your point, the more, from my perspective in the US anyway, the more non-US content we're able to get and enjoy here, beyond the enjoyment, it does open up, you know, open your mind up to other cultures and see how other people in other parts of the world do things, which is hugely important. What are some of the other things that you either might be working on or just might be thinking about going forward that breaking down language barriers like this and being able to export content in a format that I think is more natural to consume? And then And perhaps, you know, it's also able to convey the original intent a little bit better than just subtitles or just overdubs might do.
21:47What are some of the applications that, you know, you might be thinking about or even working on that listeners might not be aware of? Yeah, so it's very interesting. There are a couple of use cases that we're working with studios. For example, working on creation of voices. for example creating diversity of voices when you have a very creative actor and he wants to do several parts right we're enabling him to do several parts because he's very creative he has his ideas of how to convey this or if there is a director they know exactly how to say this piece but he wants to create it in the specific way you know and there are some directors so we're enabling them to do this with the technology.
22:34Another use case is doing a screen test. For example, currently only screen test for a new movie or a new TV series, testing the jokes, testing the plot is only done in the US. And then all the world is just, you know, if it works, it works. If it fails, a bummer. You know, you lost a lot of money. Why is that? because it costs a lot of money and it takes a lot of time and you are doing it at the first stages of creating the movie or creating the show. Without technology, you can do it very fast. So you can do it across four continents and then you can feel, get a feel if your content will work globally and not just in the US.
23:18And this is only an entertainment, but when you go to advertisement, you can make an advertisement that will work for example you take latin america so with different words in mexico and in argentina and currently i don't know if the audience know but in movies they use latin american spanish it's a natural spanish actually if you go and talk in the street in argentina and in mexico it's different spanish in a sense right but in movies they just flatten it to in terms of costs you want to do it in one time with our technology yeah definitely so with our technology they can do a version a mexican version and an argentinian version a bolivian version different version and you can use the actual language and the actual jobs that will work in each region so so and this is only in entertainment and you can go into advertisement and even if you go down the line, and I think this is the most impactful that this is our goal, e-learning, edutainment, enabling people from places where dubbing is not cost-effective to access knowledge.
24:36So we, you know, people that understand English, they have access to a lot of knowledge. But if you don't understand English, a bummer you cannot you know you cannot access most of the the the e-learning content that is out there most of the youtubes are in english that you can learn from popular science and even to learn about technology you know part of the triggering to build this company was that my brother worked in a company in brazil and he were worked in cyber security and he has people from Brazil and he wanted them to learn cybersecurity deeper. And he told them just go on this platform and learn it.
25:18And he understood they cannot do it because it's very difficult to hear something in English because you are concentrating on translating instead of learning. Instead of learning, yeah. Along the way, I learned that there are good people that create fantastic content, but unfortunately it's not in English. They have like millions of subscribers in their own language and nobody else can consume it. So for example, I've come across a French guy that has like a popular science kind of a YouTube channel. It's very interesting to learn that his content cannot reach other audiences. Right. So in an in an essence, technology can enable people that create good content, be connected with people that want to consume it, but they want to consume it in their own language.
26:15Sure. Yeah. And if you go to children, you know, children don't know other languages. So actually you want to access them early age and enable them to get to very sophisticated kind of content or knowledge because knowledge is now in the Internet. And we need also to understand, and this is something very important, that the youth currently, you know, we used to read a lot. I used to read a lot of books. I'm on Facebook. But my kids, they are on TikTok and YouTube Shorts and Instagram. They are not on the word side of it. They are on the audiovisual side of it. Yeah, when my kids want to learn something.
27:01It's actually funny because I noticed it first when we would be talking about something and we'd want to look something up. And I would go to a search. I'd go to Google and I'd type. And they'd kind of look at me like, why aren't you going to YouTube? Like, that's where you learn things. You go to video, search, and you watch a video. Yeah. So your point is very well taken. Absolutely. Right. And if you are a company, you know, I don't know what you are doing, but when I go to the website and I search for a product and I go to the website and I see a product tour, a video of the product tour, this is the first place I would press.
27:38You know, I want to see it in one minute, just understand what it's doing. Instead of reading all these words, nobody has time. And we need to understand sub-dozen words. in the hectic times that we have, you know, I'm hearing on my phone something. I have like, I have a kid that has like three screens. He has his, you know, the phone, the tablet, and the computer, and the television is all on. And I am amazed how they can split their attention, but they are doing this. You cannot do it with subs. Right. Yep. No, you can't. So looking ahead, we're recording this just about in the middle of the year, in mid-late June here of 2023.
28:21But where do you see this all headed in the next couple of years? Is it a matter of just refining the technology and being able to build usage and then bring costs down to make the tools more accessible, as you said, not just to the super high-end projects with the studios, but other users as well? Is there something else that you're working on? Where is all of this headed, you know, in whatever the timeframe is that makes sense to talk about? Yeah, so in the near future, we're working to democratize and enable people to access the technology in a very affordable manner. And in a way, they can, you know, have their content be localized and accessible.
29:08That's accessible. And this would be a huge impact because even people that don't have resources can do this kind of change or reach more audiences in an essence. But in the future, we'll also enable a real-time translation. So the technology will enable you to real-time translate an audiovisual content. So you have more news and more sports that will be available. That's what I was wondering about, yeah. Yeah, it's a matter of compute strength and advancing the translation part. But as this will go, you know, this technology will be available. I think that with our knowledge and understanding of content, which is not just creating a dialogue.
30:01There is a lot of companies creating dialogues. It's not just creating dialogue. It's just handling the end product, which is the actual content. I believe that the technology will enable people to enjoy content that is created around the world. And it will really globalize the storytelling and knowledge of people that which currently is bound by the language barriers. Fantastic. Well, the company is DeepDub, deepdub.ai. for listeners who want to find out more. There's the website. Are there other places, social media, blog, other places that you would send listeners to to find out more about what you're doing?
30:44Yes, so we have a blog on our website. We also have a newsletter that you can unlist and you can also meet us on the social media or a LinkedIn page or Twitter and Instagram page. So you can find them and register and just follow our company page and you'll get to be the first, you know, the new advancement that we are preparing for you. Fantastic. Ophir, it was a pleasure. And it's really incredible stuff you're working on. And, you know, we just kind of scratched the surface of the possibilities. But yeah, anything that knocks down barriers and helps people understand where each other are coming from and shared experiences is a win in my book.
Read the full transcript
From the publisher
In the global entertainment landscape, TV show and film production stretches far beyond Hollywood or Bollywood — it's a worldwide phenomenon.
However, while streaming platforms have broadened the reach of content, dubbing and translation technology still has plenty of room for growth.
Deepdub acts as a digital bridge, providing access to content by using generative AI to break down language and cultural barriers.
On the latest episode of NVIDIA’s AI Podcast, host Noah Kravitz spoke with the Israel-based startup’s co-founder and CEO, Ofir Krakowski. Deepdub uses AI-driven dubbing to help entertainment companies boost efficiency and cut costs while increasing accessibility.
The company is a member of NVIDIA Inception, a free program that offers startups go-to-market support, expertise and technological assistance.
Traditional dubbing is slow, costly and often missing the mark, Krakowski says. Current technology struggles with the subtleties of language, leaving jokes, idioms or jargon lost in translation.
Deepdub offers a web-based platform that enables people to interact with sophisticated AI models to handle each part of the translation and dubbing process efficiently. It translates the text, generates a voice and mixes it into the original music and audio effects.
But as Krakowkski points out, even the best AI models make mistakes, so the platform involves a human touchpoint to verify translations and ensure that generated voices sound natural and capture the right emotion.
Deepdub is also working on matching lip movements to dubbed voices.
Ultimately, Krakowski hopes to free the world from the restrictions placed by language barriers.
“I believe that the technology will enable people to enjoy the content that is created around the world,” he said. “It will globalize storytelling and knowledge, which are currently bound by language barriers.”
https://blogs.nvidia.com/blog/2023/08/30/deepdub/




