#265 Jeff Lu: How Akool Uses Gen AI to Make Live AI Avatars & Voices

27 Jun 2025 · 45 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. - Episode #265 Summary

Episode Title

Jeff Lu: How Akool Uses Gen AI to Make Live AI Avatars & Voices

Key Takeaways

  • Introduction of Akool: Jeff Liu, founder and CEO, discusses how Akool is leveraging generative AI to create live AI avatars and real-time video translation, effectively bridging language barriers.
  • Technology Overview: The episode delves into the technical architecture behind Akool's offerings, including the use of large language models for translation and the complexities of live video processing.
  • Future Implications: The conversation explores the potential future where the need to learn languages may diminish due to advancements in AI translation technology.

Episode Breakdown

  1. Creating Digital Clones with AI (00:00)
  2. Jeff Liu introduces the concept of creating digital avatars for various applications like meetings, podcasts, and streaming.
  1. Jeff Liu's Background (01:20)
  2. Liu shares his journey from working at major tech companies (Apple, Google, Stanford) to founding Akool.
  1. What Akool Does and Its Market (02:54)
  2. Akool focuses on AI video technology, including avatars for training, live streaming, and video translation. The main customer base includes businesses, studios, and marketing agencies.
  1. Live AI Avatar Suite Overview (07:11)
  2. Discussion on the suite of products offered by Akool, including different types of avatars and their functionalities.
  1. Real-Time AI Avatars (10:32)
  2. Explanation of how Akool’s technology allows for the creation and use of live avatars in real time.
  1. Language Translation Solutions (16:05)
  2. Akool’s approach to language translation through avatars and how it simplifies communication across different languages.
  1. YouTube Video Translation (21:40)
  2. The process for translating YouTube videos with Akool’s technology, including lip-sync capabilities.
  1. Technical Architecture Behind Translation (28:36)
  2. Insight into the technology powering real-time translation and the challenges involved.
  1. Competition with Tech Giants (33:22)
  2. Liu discusses how Akool differentiates itself from larger companies like Google.
  1. Akool's Future Vision (36:26)
  2. The vision for the future of Akool and the potential impact of their technology on global communication.
  1. Types of Avatars Explained (39:36)
  2. Detailed description of the different avatar types: Instant, Studio, and Ultra.
  1. Emotional Nuance in AI (43:39)
  2. Discussion on integrating emotional expressions into voice and video for a more authentic user experience.

Discussion Highlights

  • Impact of Technology: The ability to translate content instantly and create digital clones could revolutionize communication, allowing for greater cultural exchange.
  • User Accessibility: Akool aims to provide a user-friendly experience where users simply input a YouTube URL to get translations, eliminating the need for programming knowledge or complex setups.
  • Language Learning Future: The episode contemplates a future where language learning becomes less critical as technology advances.
  • Emotional Communication: The integration of emotional cues in AI-generated voices and avatars is highlighted as a significant area of focus for the company.

Conclusion This episode of Eye on A.I. presents a compelling look into how Akool is at the forefront of generative AI technology, breaking down language barriers and redefining communication across cultures. The discussion also emphasizes the importance of emotional intelligence in AI interactions, hinting at a future where technology not only facilitates but also enriches human connections.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:04You can create a digital clone of yourself and use that as a video feed into any software like meeting software, podcast software, streaming, streaming and streaming and streaming. Our current translation is a separate feature that they do translate and animate to a real video feed, not the avatar feed. So it's actually easier for us to make the avatar do the translation rather than the real video feed. the barriers of language will be pre-trained reduced in three to five years. Lots of the video translations will happen. Lots of the real-time translations will happen. Maybe in the next generation, they don't need to learn so many languages anymore.

1:04So, Jeff, can you start by introducing yourself? Give us some of your background, where you went to school, what you studied, what you were doing before this, and then tell us about Akul. Sure. My name is Jeff, and my background is mainly around the engineering side. So I spent more than a decade working on the AI videos. So previously, I did my PhD in University of Illinois at Bernard Champagne with one of the big name professors over there. And I did a research role in Stanford working on video generation. and after that I went to Apple working on the face ID working on how to generate understanding face identities and so on and I went to Google Cloud working on video processing on the cloud especially like understanding the videos, what's happening over there and just so on and this is the first company that i started about three years ago and i i have some other startup experience with some of the friends company yeah and what what is Akul's premise.

2:47What's the product? And then I'll ask you about the tech behind it. Sure. So the company was founded three years ago, and what we work on is the mainline AI videos, especially around humans in the videos. And, for example, we do AI avatars, either for training education, which create videos, or for live streaming and adapting for meetings and so on, as well as marketing, the influencers, lots of the stuff around avatars. And another thing is video translation. We do lots of the video translation stuff. You can give a YouTube link can translate the YouTube video into any other language. Hi, I'm Eugene.

3:45This is Ecolife Camera. This is our Ecolife product. It's the first first product of the Ecolife product. Hello, my name is Eugene, and this is an Ecolife Camera demonstration, the last product of our suite Ecolife, one of the first solutions in the world for video generation in real time. And also other things, including face swap, we can change the face and appearance and the age of the person in the video. There are lots more. Our main customer base are more on business side. I see. Studios, agencies, marketing companies and so on. Yeah. How many of these, I mean, you mentioned a suite of products there.

4:35How many have been launched to date? Yeah. So we already launched many products. I think we have like close to 10 products on our website that's ready to use. Most of them are video creation sites. Yeah. And does that include, you were saying you can translate YouTube videos into other languages. Does that include lip syncing the videos? Yes, that includes lip syncing the video as well. So especially if the person like doing podcasts with that kind of thing, we can do lip syncing. And we can also do more complex situations, including movings as well. Yeah. Well, on the lip syncing, I'm curious because the translate, you know, I've done, I had a company called LipDub on the podcast.

5:30and they do lip syncing of dubbed videos, or you can upload Chinese, for example, audio, Chinese translated audio to a video in which I'm speaking English. But the length of the sentences differs, So you have to do some adjustments, maybe make the Chinese sentence longer or shorter in order to match the video. And then they do the lip sync. How do you handle that with your YouTube translation? Yes. So that's a good question. We provided two options. The first option is keep the original video length. In that case, we will generate the sentence that's the same length when you speak it out in different languages.

6:36And another option is called dynamic video length. So we will change the video a little bit, either speed up a little bit or slow down a little bit. So you don't necessarily need to keep the lens to be exactly set. Oh, that's interesting. Yeah. Well, I'll have to try that. But the product that we're talking about today is this live AI avatar. Can you describe what that product is? Sure. So yeah, the products we are launching is live suites and it includes a series of products and the main one is live AI avatar. What it does is you can create a digital clone of yourself and use that as a video feed into any softwares like meeting softwares, podcast softwares, streaming, and so on.

7:48And we have additional offerings in the whole suite as well, including real-time live video translations, real-time live face swap, and we also have an upcoming real-time live AI video generation. Yeah. Now, on the live avatar, I understand training in avatar, but how do you have the avatar lip sync with the voice in streaming video, live streaming video? To me, that sounds impossible. Yes, it's a very hard task to lip sync in real time into the meetings with the voice. And we did a lot of optimization to ensure both the quality and the speed for the amateurs and the lip sync. So actually we have a series of models that run some different, uh, computation resources.

8:57On the high end, our models run on like, uh, G H 200 or H 200 GPUs. And it goes to like a 100 and a 10s. And our most optimized model even runs on the laptop GPUs. So on the laptop side, we are working mainly with AMD. And for the live avatar side, we are running on various types of GPUs based on their situation and needs. And they've been working very well. Yeah. So for people, because this is both an audio and video podcast, for people that are listening only, I'd encourage you to go to YouTube and watch the YouTube version of the podcast. Because, Jeff, what's on screen is an avatar. Can you switch to your live view, your real self?

10:03And then we can talk more about how that avatar is working in live streaming. Sure. There we go. So can you just walk through the architecture? First of all, is this proprietary, the models that you're running in the background, or is it an architecture of API calls, or how is it happening in the background? Yeah, yeah. So we developed the model ourselves, the host app, from algorithm to the data collection, training, and the deployment, data book, sustainability. So all of them myself developed. And we deployed these models on the cloud. And it depends on different situations. Sometimes the model is with larger compute, such as H2O.

11:02And sometimes it's a smaller compute. It really depends on the need. And how it works is on our desktop site, you install a small install package, the OpenEdge, and it will connect your camera with our cloud servers, and your camera stream will be streamed to the cloud servers, and we want the computer on the cloud, and that stream back to the server that you are using, such as Google Meetings or Podcaster software. Yeah. And so the video streams to the cloud, and the cloud streams back the avatar matched to the voice. Is that right? That's right. So the lip sync and avatar animation, and these things, depends on the cloud, and it's going back to the presenting software.

12:06Yeah, and do you need other, I mean, you download this package, but do you need a special software on your, I mean, special hardware, a laptop with special hardware? Does this work on consumer laptops? It's consumer laptop, so the software just do very minimal thing on laptop. It only provide user interface for you to make action and well, and take your video stream from the camera and the streaming to the cloud. So that there is no compute happening on your laptop or Mac. And so it's a very minimum computer or software requirement. And how do you deal with the latency? I noticed with the avatar, you were speaking a little more slowly, a little more deliberately.

13:10Is that because you need to allow for the latency streaming to the cloud and back? Yeah. So for most of the live features, we optimize the latency a lot. So the thing that we noticed is that for the live interactions, usually latency below half a second cannot be perceived. And below two seconds, then it's fine. And for the features we are working on, we can show and optimize the latency to be one to two seconds. And in the best case, you can go down to at a lower half a second. But most of the time is one to two seconds latency. So that's very acceptable. And there are lots of things that we can do to optimize latency and get into more comfortable real-time experience.

14:15But when you're speaking using the avatar, do you speak more slowly or is that an artifact of the system? I think that maybe I speak a little bit slower. So there should be no artifacts for the voice because the voice is almost unmodified for the source. So the main thing that we modify in the avatar is the thing. So I don't think the voice will be slowed or so on. Yeah. Can you switch to another avatar just so we can see how it looks again now that we understand a little bit more what we're looking at? Sure, sure. Yeah, let me use another avatar to join.

15:22Okay. Yeah, that definitely does not look like you. Okay, so this avatar, is it one that you trained? Is it an off-the-shelf avatar? What is it? Yeah, it's an off-the-shelf avatar that we provide on the website. So I think everyone can use it. And, yeah, it's an off-the-shelf option. Wow. And you said that you're eventually going to have live translation, audio translation, married or merged with the avatar. Is that right? Yes, that's right. So our current translation is a separate feature that they do translate and animate your real video feed, not the avatar feed. So it's actually easier for us to make the avatar do the translation rather than the real video feed.

16:29We'll make it happen. Yeah, well, it's easier because it's easier to sync the lips of the avatar than it would be a live video. Is that right? That's right. That's right. So think the lips of an avatar is easier than think the lips of a live video, especially if they're in the random people show up in the video feed and we need to see the lips perfectly. So that's definitely a harder task than you already have an avatar with you. Yeah. And this, you know, there's things are moving so quickly. Google just announced live audio translation, or I don't know if it's live translation, but where you can switch between languages while streaming, I believe.

17:36Is this a problem that's now been solved? Are you using any other models to do this? Or can you talk about the tech behind the streaming translation? Sure. Yeah, I'm happy to talk about that. So the streaming translation is not a fully mature area. and just only just for the voice side only. So if you think about it, there are so many languages out there and how much latency you get. It's a very challenging task, especially if you want to do real-time translation. So you need to balance between latency and accuracy. The more of a sentence your voice model hears before it translates, the more accurate the translation will be.

18:40So definitely if more predictions and AIs are being used, then less latency and more accuracy. And for our case, for the translation part, we are integrating with Microsoft to do the translation and also So there are some other players we are working with as well. And for the other piece, which is lip-sync and the movement, that's also very challenging. And that's our core expertise for the live translation, which makes you able to feel like you are speaking in another language. Yeah. I mean, it's remarkable because, again, as I say, I've been translating content from English into Chinese, a video from English into Chinese.

19:42And the process I'm going through is I go to 11 labs. I translate the audio into Chinese, for example. but I have to adjust the sentence lengths. And then I export those audio files, go into LipDub, train the LipDub on the video so that it can reproduce the lower part of the face and the lips, and then upload the audio and then render or generate the new video. And it's a fairly cumbersome process, but you're talking about doing that from a YouTube video. Do you just put the URL of the YouTube video into your system or do you have to upload the video to your system directly and then train the system on the speakers.

20:52And I mean, how does all of that work? And why don't we switch back to your live view? Just it's a little creepy talking this way.

21:03Sure, yeah, that's squeegee back. But do you understand my question? how, I mean, it just amazes me, the computation required. First of all, if it's a YouTube video, you need to train your model on the face in the video. Then you need to translate the audio, and then you need to regenerate the face So the lip syncs with the audio. Is that right? So actually, we have a much simpler pipeline, and the audio works well. So our product is all for easier usage. So if you want to translate a YouTube video, you only need to copy the URL of the YouTube video and paste it to the box, and choose the language you want to translate to, and click Translate.

22:06and all the magic happens in the backend. So we do everything, including downloading the YouTube video. And so we have a model that don't need to train per person. It's a generic model. Any person come in, it immediately works. So no training is needed, actually. And so we can do lip sync and we clone the voice and every magic happens behind the hood. And it's really fast as well. And if you want to have more fine tunes of the translation, you can click into the script, either before translation or after translation, click into the script, send it there, make some edit, and we can generate another translation version that matches your updates or modification of the script.

23:01But I think for most of the case, the translation is very good. So we did a lot of the translation of YouTube videos, and the translation is, I think it's almost identical to the default YouTube subtitle translation. So it's almost the same. So I think it's already pretty good translation quality, even without any modification. Yeah. What fascinates me about this is it's not geography that separates people, it's language. And, you know, I spent a lot of my life in China. I'm assuming you're from China. And the ability to share content or view content across languages, that barrier is disappearing.

23:57And I think that's a tremendous benefit to humanity because there's greater understanding of culture and greater understanding of attitudes and opinions. When you do this to YouTube videos, are you then uploading them to somewhere in China? Is there a site that is comparable? My videos, for example, I upload to Bilibili. Yeah, yeah. For my case, it's Moon and Lee. So my wife is from South Korea. She speaks Korean and her family will speak Korean. I don't understand Korean. And also we have family members from other places speak other languages as well. So it's very meaningful for me to translate the content and understand them.

24:55Like some of the videos my wife is watching, I have no idea what that is, but I can translate it and I can watch it. So I have to consume lots of the videos in my book these days. So, and I just see some good YouTube videos, maybe in another language, especially career or something i just didn't translate it and i watched it myself i didn't post it somewhere so i just watched it myself and we do have some of the videos um trying to trying to like another language on socials and so on effort is ongoing i think more on the company side not on my personal side so on my personal side most of the video i translated i just watched it myself yeah what is the most popular video platform in china today for for this kind of content yeah so the uh the comparison to youtube i think for the long long phone and it's bidi bidi it's very popular and for a short version, TikTok, right?

26:14So TikTok is the most popular for short video. And I actually watch YouTube much more. So I mainly just use YouTube. YouTube has everything. Yeah. This tech, when did, because I think I told you when we spoke earlier, I did a podcast in maybe 2018 with a researcher at Baidu, and he was trying to do speech-to-text translation because there was such a big lag time in speech-to-text. You know, the problem that he was dealing with is Baidu has a lot of English-speaking employees that don't speak Chinese, and their executives at a conference would be speaking, and there would be, you know, translated text, you know, on the side or on the video.

27:19but the translated text was always two or three sentences behind the speaker. So if the speaker told a joke, the Chinese audience would laugh, and then the English speakers would have to wait watching the text to get the joke. And by then, you know, the speaker had moved on. And the problem that this researcher was having is in the prediction. It was just taking too much time to predict the next word to close the latency gap. And particularly between a language like English and German, we spoke about that with the sentence structure being entirely different. When did that hurdle, when was that crossed?

28:19Is that purely because of the transformer algorithm? How did the technology get over that hurdle of the latency in the prediction? Yes, that's actually a very good question. So, it's always hard to battle with the latency in translation, especially all the languages are different, right? Like Korean, the structure is very different, English, Chinese, and so on. Maybe for some language, you need to listen to until the end before you know what the people actually mean. So it depends on their language, actually. So for some of the language, the latency can already be optimized very low. And also it's a battle between accuracy and the latency.

29:20And also if the people speak a short sentence more, and latency can be even lower. But the biggest improvement in this space is indeed the transformers and these edge language models and so on, they are much better at predicting and the translations and the translation accuracies and reduced latency. So, and for the most specific details, I think, and onto the expert here, we are actually using some of their third-party services for the live real-time translation and how to how to convert a voice language from one language to another so we are integrating some third parties. They definitely did a lot of optimizations to reduce the latency for the prediction and so on.

30:24They also provide us parameters to balance latency and accuracy. It's something that we can play with. Yeah. And so does the latency change depending on the language pair? I mean, for example, Chinese and English are very similar in sentence structure. But as you said, Korean and Chinese or German and Chinese, it's quite different. Yeah, I think it depends on the language pair as well. So Chinese and English, the words order, pretty similar. And I think for Korean, maybe German, the verb is usually end on the sentence. So it's a different structure. So a word pair is highly related to the latency.

31:14I see. So in English to German, the latency would be higher. And so the video, if I were doing a live video, there would be a lag. Yeah, that's right. So basically, for some other video, you need to listen more before you can translate more accurately. You can also force it to translate. But in that case, then the translation will be not accurate. But in general, we already optimized latency to be pretty low. So definitely more and more optimization. It's nice and also language. So Akul is the name of the company or the name of the product? So it's the name of the company. And how big is your team now?

32:10Yeah, we have 70, 80 people. around and most of them are engineers. I see. And who are your backers? Yeah, the company, the mainland bootstrapped. We do have some investors, but I think mainland bootstrapped and the business is, I think it's in a place that we have pretty job product market fee. Yeah, absolutely. Are you concerned because as I said, Google has come out with the ability to translate, switch between languages and live translation, which sounds similar to what you guys are doing. Are you concerned about building market share with the big guys looking at the same space? So I think our product is not the same.

33:26Based on my understanding, what they did is they they do a voiceover on top of the original video for the translation. We are not voiceover. We actually edit the video track and also edit the video in real time to do the translation. So our solution is actually better than just doing voiceover on top of that. And also Google, usually many of these players when they introduce a solution, they bring it into the into the Google suite, right? We have a much more generic solution that you can integrate into any software. So, and also it's actually a good thing that when there are more players over there, it means this field do on top of their needs and market and wealth, as we can get much more users easily.

34:37Yeah. And how are you pricing this? Do you have a freemium model? Is it based on usage and that you can use a certain amount for free and then you have to pay? I mean, how do you do that? So the core line is currently in beta. So for this beta version, it's provided for free, but with limited access, you need a flight to get access. And for the GA launch, it will be a substitution model, and you can free try it for an hour or so. And then after that, you need to subscribe to plans and then you can do that for example 30 hours a month for them for either translation or avatars or more of the rumpster features.

35:36Yeah. And so that's for the entire suite of products or can you, for example, I'm very interested in the YouTube translation product. I have less need for the live avatar product. Do you break them out or is it bundled in a single subscription? So all of them are bundled in a single subscription that we have. So we actually have a suite of tools and offerings, and all of them are being bundled, which is a big advantage that we have all of them. Yeah. Who are the customers so far? Yeah. So we are a platform that targets more professional use and business use from the very beginning. So in the early days, we have lots of users in the marketing phase.

36:36We have many marketing agencies. We have lots of the brand and marketing department on the platform. and then we go more on to the studio and the creators and influencers. And then we have lots of technology companies on the platform and using it. So I can give some examples. Our first large client, say the Coca-Cola, they use us to run marketing campaigns and so on. And now we work very closely with tech companies like AWS, AMD, helps them to provide solutions and so on. So definitely we are across different verticals. Do you think eventually your technology will be embedded in teleconferencing software so that the end user doesn't have to subscribe?

37:38that it'll be an option in whatever teleconferencing software you use that you can, you know, choose an avatar and a language and that it takes place on the side of the teleconferencing software? Yes, that's actually a very good question. So what we believe is first we need to reduce the compute cost for lots of the things that we do. It's one of the efforts we've been working on. So if we can make most of the compute happen locally on the laptop, then I think lots of the software, they will be able to provide that as a brief feature on the software. If still a lot of computers are needed, I think it will be more likely to be a paid feature.

38:38And based on our understanding, most of the providers on this market, they are more focused on the voice piece. We didn't see too much players working on the video feed or the imaging. more on the voice side. So maybe voice will become a default option in the near future, like probably like two or three years or something. On the video side, we didn't see things happen yet. So we are still like the leaders and the early players in this market. How long does it take to train an avatar, a personal avatar for use on your platform? Yeah, so we provide all three types of avatar on our platform. So the first type is called instant avatar.

39:45For that one, no training. Upload a video, instantly works. And for the second one, it's called a studio avatar. I'll upload a video, you will check on your video. And that takes about 12 hours, but everything happens automatically on the back end. And the third one is called Ultra Avatar. So with that avatar, you can control all the body motion and hand movement, all the, give you much more control rather than just, just the face and so on. You can control the body at a moment as well. With that long, it takes longer. So our support team need to get involved to help you to create that. That might take up to three days.

40:40Yeah. And what's your vision for the future? I mean, you know, Transformers have only been around since, what, 2017, so not even 10 years. GPT-3 and onward have only been around for a few years, and you guys have only been doing this, what, three years. when you look out five years do you think the barrier between languages will will really disappear i mean what what do you envision yeah that's uh that's a goal we are working on right so definitely um we want to remove the barrier of the language and so on it's also very helpful for me personally so we're definitely working toward and then we do believe that the barriers of language will be greatly reduced in the next three to five years.

Read the full transcript

41:47Lots of the video translations will happen. Lots of the real-time translations will happen. So maybe in the next generation, they don't need to learn so many language anymore. That's right. Yeah. And this Baidu researcher that I talked to seven years ago, he was imagining an earpiece maybe connected by Bluetooth to your phone where you could have a conversation and it would be translating the other person's speech live in your ear. with a low enough latency that you could have a natural conversation. What's your thought on that? I think on the hardware side, that should be mature already. I think I already see prototypes or some kind of products.

42:52Currently, the main thing is still on the AI side, how to ensure the transmission accuracy. If you do voice clone, how to do that really well, and how to make it sound natural, and how to make it to contain the emotion in the transmission, reduce the latency. So all of these are software-related and AI-related. On the hardware side, I think that's definitely very doable for now. Yeah. And on Emotion, for example, what are you guys working on right now? What's your roadmap? Is Emotion one of them? So Emotion is a small piece in the current version, but we're working on improving the emotion in the translation and so on.

43:48So the emotion is mainly used as transfers from the original audio track to the new audio track with the emotions. That's definitely a big piece that we've been working on. And our first voice emotions on the amateurs. And the next will be trying to make the video translation to copy emotions as well. Hello, my name is Eugene and this is a demonstration of Equal Life Camera, the most recent product in our EcoLife Suite, one of the first solutions in the world for generating video in real time. This includes the broadcast of the host, the avatars of the live, and the translation of video in real time.

44:31Hello, my name is Eugene and this is a demonstration of Equal Life Camera, the last product in our EcoLife Suite, one of the first in the world solution for generating video in real time. This includes the replacement of the face in real time, the live avatars and the output of the video in real time.

45:11the invitation to get access to and that's from the days and things are coming so it's very exciting moment

From the publisher

What if you could translate any video into any language—instantly—and make it look like the speaker was really speaking it?

 

In this episode of Eye on AI, host Craig Smith sits down with Jeff Liu, founder and CEO of Akool, to explore how AI avatars and real-time video translation are eliminating global language barriers. With a background at Apple, Google, and Stanford, Jeff is leading Akool to the frontier of generative AI: cloning faces, voices, and emotions in live video streams.

 

We dive into the technical architecture behind Akool’s real-time avatars, the role of large language models in translation latency, and what the future looks like when the need to "learn languages" may disappear.

 

Whether you're a content creator, marketer, technologist, or just curious about what’s next in AI and communication, this episode is a must-listen.



Stay Updated:

Craig Smith on X:https://x.com/craigss

Eye on A.I. on X: https://x.com/EyeOn_AI

 

(00:00) Creating Digital Clones with AI

(01:20) Jeff Liu’s Journey from Big Tech to Startup Founder

(02:54) What Akool Does and Who It's For

(07:11) Inside the Live AI Avatar Suite

(10:32) How Akool Powers Real-Time AI Avatars

(16:05) Solving Language Translation with Avatars

(21:40) Translating YouTube Videos with One Click

(28:36) The Technology Behind Real-Time Translation

(33:22) Competing with Tech Giants Like Google

(36:26) Akool's Vision

(39:36) Avatar Types: Instant, Studio, and Ultra explained

(43:39) Building Emotional Nuance into Voice and Video

More from Eye On A.I.

All 266 episodes
#265 Jeff Lu: How Akool Uses Gen AI to Make Live AI Avatars & VoicesEye On A.I. · 45 min
Listen in VO