In short
No Priors Podcast Episode Summary: Can AI Replace the Camera? with Joshua Xu from HeyGen
Podcast Overview Title: No Priors: Artificial Intelligence | Technology | Startups Hosts: Elad Gil & Sarah Guo Guest: Joshua Xu, Co-founder and CEO of HeyGen Episode Title: Can AI replace the camera? Episode Release Date: [Insert Release Date] Episode Description: The episode explores the advancements in AI video generation, focusing on HeyGen's technology for creating personalized video content using language, video, and voice models, and discusses the implications for deep fakes.
Key Concepts Discussed
Introduction to HeyGen
- Joshua Xu shares the origin of HeyGen, which aims to revolutionize video content creation.
- Background: Xu previously worked at Snapchat for six years, focusing on AI-powered camera features.
Vision of Replacing the Camera
- Core Idea: AI can serve as the "new camera," reducing barriers in visual storytelling.
- Challenges in traditional content creation include scheduling, costs, and the need for professional setups, which can hinder content generation.
Breakdown of Video Production Process
- Video Elements:
- Camera (Avatar): Represents human spokespersons.
- Editing (B-Roll): Involves additional assets like voiceover and music.
- The team disassembled the video production process to identify areas for AI enhancement.
Applications of HeyGen Technology
- Three Use Cases:
- Create: Generate explainer videos and training content.
- Localize: Translate existing videos into over 175 languages.
- Personalize: Tailor video messaging for individual recipients.
Quality and Model Development
- Emphasis on maintaining high-quality standards (above 90) for video outputs to replace traditional methods effectively.
- The technology stack includes in-house video creation models, alongside OpenAI’s ChatGPT for text generation and voice modeling.
Future of AI Video Generation
- Discussion of potential real-time applications and the evolution towards personalized video content at scale.
- Real-time avatar generation could replace live interactions in customer support or presentations.
Research and Development Challenges
- Integrating various components (text, voice, video) poses unique challenges, particularly in ensuring cohesive output.
- The focus on aesthetics in video generation is crucial as customer satisfaction is a primary metric for success.
Safety Measures Against Misuse
- Strict policies against the creation of unauthorized content, including political material.
- Advanced user verification and safeguards in place to combat deep fakes and ensure responsible use of technology.
Key Takeaways
- AI Video Creation Potential: Technology can democratize content creation, making it accessible for businesses of all sizes.
- Quality Threshold: Maintaining high standards in video generation is essential for customer acceptance and market success.
- Future Innovations: Advancements in full-body avatars and real-time generation could transform industries beyond marketing, including education and customer service.
- Ethical Considerations: As AI capabilities grow, so does the responsibility to mitigate risks associated with misinformation and misuse of technology.
Conclusion The episode highlights the transformative potential of AI in video production and the ongoing efforts of HeyGen to pioneer this change. With a commitment to quality and ethical use, HeyGen stands at the forefront of a new era in visual storytelling.
Links
- [HeyGen](https://www.heygen.com/)
- [Follow on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @joshua_xu_](https://twitter.com/NoPriorsPod)
Feedback
Email feedback to
[show@no-priors.com](mailto:show@no-priors.com) Subscribe for new episodes on [Apple Podcasts](https://podcasts.apple.com), [Spotify](https://spotify.com), or wherever you listen! [Find transcripts](https://no-priors.com/) and sign up for updates.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:05Welcome, Joshua. We're so excited to have you here today. How are you? Hey, Sarah. I'm so excited to be here. Thanks for having me today. It's our pleasure. Let's get started. Welcome to the Huberman Lab podcast, where we discuss science and science-based tools for everyday life. I'm Sarah Guo, and I'm a professor of neurobiology and ophthalmology at the School of Medicine. Wait, Sarah, I'm so confused. What's going on here? Is this thing on? Today, we're here to discuss how AI can benefit your health and what medicinal properties the technology holds. Sarah, I'm so lost. Isn't this the No Priors podcast where you interview technology superstars like Gary Tan and Alexander Wang?
0:46No, that's only for humans. We're really excited to have you. Welcome, Joshua. Yeah, excited to be here. Thank you for having me. So let's start with a little bit of backstory. You started this company, Heijan. It's had this amazing growth trajectory and is being used by millions of people now. What's the story of starting the company. Yeah, sure. So yeah, hello everyone. My name is Joshua, co-founder and CEO of H &M. We founded the company roughly three and a half years ago. And before that, I was working at Snapchat for about six and a half years there. I started robotics at Carnegie Mellon and joined Snap back in 2014 there.
1:25I initially worked on machine learning in Snapchat ads, ads ranking and recommendation than I spent my last two years at Snap working on AI cameras. So, you know, Snap leveraged a lot of AI technology to enhance the camera experience. If you look at, you know, 2018 Snapchat released a baby filter and Disney style filter. That was the first time I saw a computer can actually create and generate something that does not exist in the world. I was just so fascinated by the technology back then, and I had a feeling that that would potentially change the with how people create the content. So, you know, Snapchat is a camera company and everybody created content through the mobile camera.
2:06But we wanted to replace the camera because we think AI can create the content and AI could become the new camera. And that's how we get started with HeyGen and our mission is to making visual storytelling accessible to all. I love it. The greatest minds of our generation, you know, inspired by, you know, your face is a cute kitten or whatever. What does replacing the camera mean to you? Like, why do we need to do that? I use my camera a lot. I kind of grew up my career in the whole mobile camera space where we work on a lot of the software and technology to enable people to feel comfortable and make it easier for people to create content through the mobile camera.
2:48But there's still lots of people are not able to create good content using a camera today. and we felt that if we can replace the camera, that means we can remove the barrier for visual storytelling, for visual content creation and that will help us to step ahead in terms of the whole content creation space. What are some of the areas that you think, you know, the technology that you developed has applied to? Because I think you've started with different forms of like virtual avatars so that you can take a video of yourself and then turn it into an avatar that you can then feed text to. It can speak in your voice.
3:24It can do all sorts of really interesting things for different areas. How did you both decide to start with avatars? And then where do you think the main applications are? When we initially started the company, we tried to like disassemble the whole video production process. It's really about camera and then editing. So camera is more about a role, which represent the human spokesperson, the avatar piece. Editing is more about B-roll, adding, you know, different assets, voiceover, music, transition, animation, stuff like that. So editing, we just learned from customers that editing is not that expensive because it's pretty standard service.
3:59But camera is super expensive. And imagine, you know, it's a CEO of a company. He wants to record something. We probably need to schedule that ahead of, you know, two weeks of time. We need to bring in the camera crew, have a studio to actually record it. And even for two minutes of footage, sometimes we need to record it for 20 minutes because people need to remember the script. And that's the piece that blocking a lot of a business create new content. So that's how we get started from, you know, trying to replace that piece of the process and making avatar to replace a camera for the video production.
4:33Where do you think that goes in the future? So, you know, people are already using HN for all sorts of different application areas in terms of, you know, marketing and sales and in some cases like internal webinars or learning or other things. I'm a little bit curious, like, is the eventual form of this, you know, everybody has somebody who steps in for them for their Zooms or is it used for entertainment purposes? Or how do you kind of view the evolution of this sort of technology over time? Yeah, I would say there's many possibilities out there. I think what we are tackling the problem so far, it is the, you know, entry point of the content creation where all the content started with the camera.
5:14And then we would have people doing a lot of editing after that. We can clearly see a path where people can already assemble all this generative footage and apply the AI editing to assemble the final video. And again, if we push forward into technology, making the performance much better, I think we will be able to create experience like generative video in a streaming way. And that actually will potentially replace a lot of the real-time conversation we have today, especially with the GBT4.0 and with all this multi-model real-time streaming technology altogether. Okay. We're still in asynchronous video creation land in 2024.
6:00How do people use HeyGen today? What are the favorite use cases you have? I would categorize the use case of HeyJet into three. Create, localize, and personalize. And, you know, people can select, you know, the cast from a library from our avatar or create their own digital twin and just like select a template or type the script and generate a video. This works the best for product explainer, how-to videos, learning development, and some sales enablement training content. We can also take existing video that localize that into more than 175 different languages and dialysis. And in this way, we can help customers to really localize their content into local languages.
6:48And last but not least, people can also use H &M to personalize the video messaging at scale. So I think there is many, many very creative use case on H &M today. We are a very horizontal platform. I would say one of my favorite use case is probably the recent launch with Madonna. And Madonna launched a sweet campaign where they can allow people to send a message to a family member in the different languages. I love him so much. I would run out of words to express my love for him. You know, I just want to call it out, you know, AI is for everyone, grandma and grandchildren alike. Yeah, that's really cool.
7:27I mean, that's a big brand in a public consumer-facing use case. How do you think about the quality of Heijen today? I would have thought of that as tip of the pyramid in terms of quality. And how can you tell when the avatars are good enough and not? Yeah, so I would say quality has always been the number one priority of the product and business and technology. I would say, you know, I always have frameworks like this. There's an invisible line of quality. You know, let's say that threshold is 90. Anything below 90 essentially is unusable for the customers because we cannot really replace the real-life production process they have.
8:16We really need to focus on making the video generation quality that go above that threshold. And I think especially for avatar today, it is above that. So we can really helping people to replace the real camera and unleash a lot of our creativity process that help people to scale the content production there. And, you know, obviously, you know, there's much more room to improve, for example, generating the full-body avatar, being able to bring all source of our element into a video. Yeah, we're in the process of that. What are you most excited about in terms of, like, what's next or new releases you guys have coming?
8:55I think there's many things very excited going on in our technology and product roadmap. I think particularly I'm very excited for the full-body generation of the avatar. Historically, all the other technology has been focused on the upper body. It's really hard to generate the gesture and the body motion, but a lot of academic research has proven that this is very possible now, and we just need to basically take that into the last mile. And another thing I would say, something I'm very excited about, the streaming avatar, especially with the latest release on GBD4.0, really, really help to improve the performance of the real-time interaction with text and voice.
9:36And H &M avatar could become a visualization layer for all those applications. Obviously, you need full gesture control and movement to get to any video of any kind. But what do customers want to do in terms of full body motion today? You had a demo of walking in the last couple months. The way how we look at it is that there is a spectrum of the quality requirement laying out on different use cases. Let's start from the left side of the spectrum. It's the learning development content, educational content. It's more like one-to-many broadcasting, talking about educational training content. the quality there is lower because the avatar can be more still more professional but if on the right side of the spectrum we call that as like the high-end you know marketing content really dynamic you know the the one example would be the ads creative and people ship very very dynamic content on ads and because that can really help to improve the ROI of the content making it more engaging.
10:47I think making the full body, enabling that full body rendering will be able to help us to bring the avatar, to bring the video into the next level of engaging and authentic. And that will help to unlock a lot of use cases in a broader case of marketing and sales. Newscasts or other things, to your point, they often have the shot of the people walking and talking as like a standard canned shot. And there's like these standard things that they use that if you had full body, you could provide for all sorts of application areas. I guess related to that, what is the technology that you folks are using today?
11:18You mentioned some things like GPT-4-0, but you've also built your own models in-house. How do you think about the technology stack that you're using, and how does that have to evolve in order to be able to do full body or other new things? There's three models, right? Text, voice, and video. So we work with OpenAI, ChatGPT. on the text generation side, obviously also serves like the brain of the orchestration engine that we build internally. And we work with, you know, OpenAI and Event Lab on the voice engine, but we build the entire video stack in-house, including after creation, video rendering, and B-roll generation.
11:56So I think over time, I think the whole technology trend has been moving towards to a direction. A lot of all these things will be chained together. The multi-model model, multimedia, or get into one single model. One of the challenges I want to call out for the full body generation is actually how do you actually connect that voice together with gesture motion? And that's actually something will be unlocked by actually getting the voice model and the video model training together so that it can sort of build a connection underlying the model as well. And that has been historically really, really hard because we have to train the TTS model on one hand and then feed the TTS model outcome into a video model.
12:43And it's pretty hard to build that connection. But with multi-model model training, that's very possible. Obviously, Sora is not available to developers and end users today, but there are world-class text-to-video generation models that are generic, not avatars. How does this technology differ from something like Sora? When we initially started HeyGen, we want to help the business solve the video creation problem. What is the business looking for? They're looking for quality. They're looking for control. They're looking for consistency, right? So when we try to look back, okay, this is not that.
13:19How can we get there? What's the technical path to get us there? This is essentially probably potential two paths. One is the test to image, the server, where we try to generate the entire thing from end to end. and so you get an entire video at once. And the other approach is that what we believe in at Hadrian is that we try to assemble the whole video into different components. Largely, it will be A-Roll and B-Roll. B-Roll represents all different kinds of elements like voiceover, music, transition, A-Roll being the avatar. And we try to tackle this component one by one and then we build an orchestration engine around that to assemble that final video together.
14:01We felt that this technical path is more capable to deliver the quality, you know, the control and the consistency that the brand is looking for. Because, for example, there's some stuff we probably should not try to generate. It's the logo and the fonts. That needs to be very accurate. And not to mention that we also need to be able to learn about, especially in the business context, we need to learn about the brand style, the color mapping, essential from our customers. And I think the second approach would give us more flexibility and capability to build that system around it. And in fact, we actually see Sora as our partner because we are able to integrate that as one of the component generator and then feed that into our acquisition engine for the business application.
14:49How do you think about what research, you know, if you just focus on like components of the experience, in particular the video stack being the thing that you really want to own and be state of the art in at Heijen, how do you approach like new capabilities from a research perspective? Is it, you know, look at what's available in academia, look at the problems customers give you sort of de novo? I would say it's a combination. I would add one more thing is that we need deeply understand the limitation around the model and try to find the connection between what is the customer looking for, what's capable with the technology.
15:28Like when we really try to look at it, all the AI model had some sort of a limitation. and I think the key question is that in order to deliver a great product experience for the customer is that how do we design a product around it so that we can try to avoid the limitation of the model but help to amplify the strength of the model and this is something that is really important to find a new area that unlocked a new creation experience. One example would be when we look at the video translation technology is a whole new way to translate your content compared to traditional dubbing. We preserve the user, the natural voice, and their facial expression.
16:13But if you look at really underlying the model, what enabled that video rendering, it is actually a lip sync model, right? But we kind of have figured out a way to combine all this together, together with the voice, as well as the translation with ChatGPT and build a great experience around it. And sort of like we are creating a whole new experience for localized video and content. So there's lots of like great McDonald's like exciting commercial applications. I think a lot of people also think deep fakes are really scary. Like and the ability to, you know, abuse somebody's likeness or voice is scary.
16:52How do you think about safety, election safety, abuse? First of all, we do not allow any political or election content on our platform today. HHS policy strictly prohibit those creation of unauthorized content, and we take abuse of the platform seriously. So we have our safety, security safeguard, include very advanced user verification, include live video consent, dynamic verbal passcode, and rapid human review in the back of all the other have been created on the platform. Trust and safety is critical to our business, and we are actively partnering across the industry, you know, continue developing the tools and best practices to combat misinformation and AI safety.
17:38And we actually build the safety as part of the design. If you look at a lot of, you know, after creation process on H &M, and we pay all this safety concern and safety guard on every single step of the creation process as well. It makes a lot of sense, I guess. It's kind of interesting because if you think about it, at least from the positive version of this, and you talked about how you try to protect against the negative, the positive version is, you know, you're running for office and you should be able to send a personalized message to each voter literally into their inbox with a short video clip of you talking to them specifically or talking to issues that they specifically care about or things like that.
18:16And so you could imagine using this technology in the future for actually hypopersonalized political campaigning. And as long as you can avoid some of the deep fake side of it, then obviously it could actually be quite valuable. How do you think this ability to really generate large scale, differentiated, personalized, et cetera, content of individuals talking, how does this kind of generation change how people make or use video in general? You know, if people can generate very engaging and authentic video content, they will basically create more videos and use video more for their business to grow their business.
18:58And we live in a video first row where every business wants to create more videos. I think the bottleneck today in the industry is that video is just very expensive to make. And it takes like weeks or months to make a video. I think it would fundamentally change a lot of ways how people are thinking about how to grow the business, how to do the communication, how to do the marketing and sales. So I do think there's a huge possibility that we can create and generate a very high degree of personalized video, especially with the full body avatar that being able to deliver a very dynamic and high quality content out there.
19:43So I want to give you one example is that I think a lot of AI generation is not only about, obviously, you know, cost saving and time saving is one aspect of the valid prop. But it's actually we are seeing a lot of customer using that to unlock new use cases. And being able to do something they were not able to do before, I think that's the key job input of a lot of a business outcome today. How do you think about it in the context of real-time versus asynchronous? It feels like a lot of these technologies are focused right now on asynchronous use cases. And that's true as well of just pure text-to-speech models.
20:18When do you think we move to any sort of real-time or close to real-time video avatars and sort of the uses of that? I look at it in two ways. One is that the real-time application of the avatar, even now it's possible. I think people can already experience that on H &M. We are making new updates that can make it even faster. so it can potentially become, let's say, the virtual AI, SDR, virtual support that help to take customer calls or provide support.
20:53And I think the technology has been always developing like this trend. Two years from now, it would not be crazy to look at a lot of avatar generation. Asynchronization pipeline will become real-time streaming capable. And I also see the world is moving towards a way that we can probably generate the entire video in real time as well in the future, let's say five years from now. I have an opinion like, you know, generative image is still image, but generative video is not a video. It is a new format. What I mean by that is, you know, when we really look at video, we look at it as an MP4 file, right?
21:34So it is immutable. For example, if you and I are on Instagram, we probably get recommended by two different ads. But as long as we are recommended from the same business, we are looking at the same MP4 file. But it does not need to be the same. Let's say if maybe I like avocado, I should be watching an ad with Coca-Cola and avocado and showing the new story about Coca-Cola with me. And you like something else, you could be looking at something else. And this is not possible today because making a video is expensive. But this could be very possible. Let's say we can actually, you know, real-time generating the video ads that you like according to your user attribute.
22:18That will potentially become a new format. You know, when we really look at today's video player, it corresponded to only one MP4 file. And it doesn't need to be true. It doesn't need to be like that. that video player can actually take in a lot of, you know, user attributes and generate something in real time to match what's the best way to deliver their content to customers. Yeah. Yeah, I think it's, you know, one interesting analogy would just be like if you think about, you know, YouTube as one of the largest learning devices in the world today. Like it is static, immutable video for everyone.
23:01But it's pretty clear, Bloom Studies and everything else, that personalized education is going to be the path that is more effective. And people want to learn by video, but it's very hard. It's too expensive to make that video personalized. This feels like an opportunity for a very different educational future too. Yeah. And one of the use cases we have seen from customers is that PubSys Group, they generate more than 100 ,000 videos, a thank you video to send to all the employees globally and localized into different languages, personalized with a name and what they like about it, you know, when they join the company and stuff like that.
23:42And historically, that is actually only delivered with one video, right? So maybe the CEO of the executive team hop on into a camera and we call something, you know, saying thank you for the, you know, 2023. But now that message and communication can really personalize at a very, very big scale. So one thing you mentioned is the various aspects of research that you're doing in terms of building your own video models, as well as using third-party APIs. What's been difficult or hard from a research perspective? Unlike a lot of other models, I think building video models, being able to integrate aesthetics into the AI models is pretty hard.
24:20So, you know, video generation is not only about solving a mathematical problem. It's actually about creating something the customer loves and appreciates. So essentially, a model with a lower optimized cost function doesn't mean it actually produces a better visual outcome. So I guess that is the piece that making it really hard to evaluate, but also really important to deliver the last mile of the value for the customer. And, you know, generally evaluation is also hard. We have to rely on in-product signal, for example, A-B test, to know which model is actually better. Because, you know, only the customer can be the judge for that.
25:01And this process generally is just not differentiable from a mathematical standpoint. We kind of have to form a system, build a system around it, and be able to feedback those data into our model training so that we can continuously to improve. Did this approach come to you because of your work at Snapchat, working at consumer products, or is it something that you had to come up with in the context of, hey, Jan itself? I would say it's very similar, especially when we work on the camera software. So how do we know whether this parameter works better or the other one works better? And I think we can definitely come up with some very objective matches about, hey, lightning score, this is lightning score, this is resolution.
25:45But there's many things we think out, hey, a better resolution, I mean, high resolution doesn't mean it's a better image quality for customers. If you look at iPhone, it does not have the best resolution always compared to a lot of other phones, but it does produce the image that most people like about using iPhone to capture the image. And yeah, there's a very similar lessons out there we learned from in early days of Snap. Yeah. What can you say about how big Heijun is today? We are a little bit over 40 people, but we are serving over 40 ,000 paying customers on the platform today. And I think what's so interesting about our customers is that these are not the typical AI early adopters.
26:32These are mainstream companies from European manufacturers to small business to global nonprofits to Fortune 500 companies, which is the problem we are solving. Given that you have 1 ,000 customers per employee, which is an incredibly impressive metric, are there specific key roles that you're hiring for or other things that maybe members of our audience may want to apply for? Sure, yeah. We're hiring across different teams, basically. Product, design, engineering, AR research, and go-to-market. Yeah. This has been a great conversation. Thanks, Joshua. Thanks so much. Yeah, thank you. Thank you for having me.
27:07Find us on Twitter at NoPriorsPod. Subscribe to our YouTube channel If you want to see our faces, follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week. And sign up for emails or find transcripts for every episode at no-friors.com.
From the publisher
AI video generation models still have a long way to go when it comes to making compelling and complex videos but the HeyGen team are well on their way to streamlining the video creation process by using a combination of language, video, and voice models to create videos featuring personalized avatars, b-roll, and dialogue. This week on No Priors, Joshua Xu the co-founder and CEO of HeyGen, joins Sarah and Elad to discuss how the HeyGen team broke down the elements of a video and built or found models to use for each one, the commercial applications for these AI videos, and how they’re safeguarding against deep fakes.
Links from episode:
HeyGen
McDonald’s commercial
Sign up for new podcasts every week. Email feedback to show@no-priors.com
Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @joshua_xu_
Show Notes:
(0:00) Introduction
(3:08) Applications of AI content creation
(5:49) Best use cases for Hey Gen
(7:34) Building for quality in AI video generation
(11:17) The models powering HeyGen
(14:49) Research approach
(16:39) Safeguarding against deep fakes
(18:31) How AI video generation will change video creation
(24:02) Challenges in building the model
(26:29) HeyGen team and company




