We are not ready for better deepfakes

24 Jul 2025 · 57 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: "We are not ready for better deepfakes"

Podcast Title

Decoder with Nilay Patel Guest Host: Alex Heath Guest: Gaurav Misra, CEO of Captions Air Date: Thursday episode, date not specified

---

Episode Overview This episode of Decoder features a conversation between guest host Alex Heath and Gaurav Misra, the CEO of Captions, a company specializing in AI-generated videos, particularly deepfakes. The discussion centers around the evolution of deepfake technology, its implications, potential misuse, and the ethical responsibilities of creators in the industry.

Key Points of Discussion

  • Deepfake Technology Progression:
  • Generations of Deepfakes:
  • First Generation: Face swaps using real footage with real actors.
  • Second Generation: Lip-syncing real footage to create videos of people saying things they never actually said.
  • Third Generation: Full-frame generation techniques that create videos from scratch without the need for base footage.
  • Fourth Generation: Future models that will allow for unlimited video lengths and enhanced context-aware audio.
  • Current Challenges:
  • Identifying and regulating the misuse of deepfake technology.
  • Lack of legal frameworks for criminalizing harmful uses of likenesses in videos.
  • Problems with audio realism and synchronization in generated content.
  • Ethical Considerations:
  • The balance between innovation and responsibility in developing potentially harmful technologies.
  • Misuse of technology can lead to significant misinformation and societal trust issues.
  • The importance of creating awareness about deepfake technology among users to foster critical viewing skills.

Current State of Deepfakes

  • Market Presence: Deepfake technology is being widely adopted for marketing purposes, enabling small businesses to compete in a video-centric digital landscape.
  • User Responsibility: Captions encourages users to utilize their technology responsibly, while also monitoring and moderating content to mitigate misuse.

Future Outlook

  • Regulatory Needs:
  • Call for comprehensive laws to hold individuals accountable for malicious use of technology.
  • Need for technological solutions that ensure authenticity in media, such as cryptographically signed videos.
  • Potential Outcomes:
  • A future where societal trust in media could diminish, reverting to a state where individuals rely heavily on personal trust rather than factual verification.
  • Development of tools and systems that can clearly differentiate between authentic and manipulated media.

Conclusion The episode highlights the rapid development of deepfake technology and the dual-edged sword it represents, offering creative possibilities but also raising ethical concerns. Misra emphasizes the need for continued dialogue on the implications of this technology to ensure responsible use while navigating the inevitable advancement of AI.

---

Additional Resources

  • [Captions Blog Post: "We Build Synthetic Humans. Here’s What’s Keeping Us Up at Night"](https://captions.com/blog)
  • [Related Articles from The Verge](https://www.theverge.com)
  • Discussions on AI-generated video advancements and the societal impacts of deepfake technologies.

---

Credits

  • Producers: Kate Cox, Nick Statt
  • Editor: Ursa Wright
  • Theme Music: Breakmaster Cylinder

For feedback or suggestions on future episodes, listeners can reach out via email at decoder@theverge.com or connect through their social media platforms.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Support for the show comes from Public, the investing platform for those who take it seriously. On Public, you can build a multi-asset portfolio of stocks, bonds, and options, and now generated assets, which allow you to turn any idea into an investable index with AI. Go to public.com slash podcast and earn an unkept 1 % bonus when you transfer your portfolio. That's public.com slash podcast. Paid for by Public Investing Brokerage Services by Open to the Public Investing, Inc. Member FINRA and SIPC. Advisory Service by Public Advisors, LLC, SEC Registered Advisor. Generated assets is an interactive analysis tool.

0:33Output is for informational purposes only and is not an investment recommendation or advice. Complete disclosures available at public.com slash disclosures.

1:00scales with you. That's why top startups like Cursor, Linear, and Replit use Vanta to get and stay secure. Get started at vanta.com slash vox. That's v-a-n-t-a dot com slash vox. Vanta.com slash vox. Support for this show comes from Odoo. Running a business is hard enough, so why make it harder with a dozen different apps that don't talk to each other. Introducing Odoo. It's the only business software you'll ever need. It's an all-in-one, fully integrated platform that makes your work easier. CRM, accounting, inventory, e-commerce, and more. And the best part? Odoo replaces multiple expensive platforms for a fraction of the cost.

1:46That's why over thousands of businesses have made the switch. So why not you? Try Odoo for free at odoo.com. That's O-D-O-O dot com.

2:01Welcome to Decoder. I am Alex Heath, your Thursday episode guest host, deputy editor at The Verge, and author of the Command Line newsletter. If you're like me, lately you've scrolled past something on social media and thought, wait, was that real? Deepfakes are everywhere, and they're getting a lot more convincing. That brings me to my guest today, Gaurav Misra, the CEO of Captions. You may not have heard of Captions yet, but by now you've probably seen a video that was generated using its AI models. The company's Mirage Studio platform lets anyone generate AI versions of real people, and the results are alarmingly realistic.

2:38Captions just put out a blog post titled, We Build Synthetic Humans, Here's What's Keeping Us Up at Night. It's a good overview of the state of deepfakes and where they're headed. As the CEO of a company building deepfake technology, I wanted to know what specifically keeps Gaurav up at night, which you'll hear us get into. I'm generally more optimistic about the long-term impacts of AI than a lot of people, but as you'll hear in this conversation, I'm a lot more nervous about this topic. Ultimately, I came away from this episode unsettled by the fact that the deepfakes of today are the least believable they'll ever be.

3:16We are not ready, and the companies building this tech, they're racing ahead anyway. Okay, here's my conversation with Gaurav Misra, the CEO of Captions.

3:43grav i'm looking at my notes here and the first thing i have is just damn we are cooked which is the text i sent you back in response to you sending me this state of deep fake report that you all put out recently and the accompanying deep fake video of yourself, which is just crazy to see how fast this is all moving. And I feel like that's actually been a theme of our texts over the past several months is you sending me something and me just going, oh my God, I can't believe this is real or not real. And I was actually thinking back to, you sent me that video in April when I was covering the meta antitrust trial in Washington, D.C.

4:26for The Verge. And you asked me, of course, thankfully, first, but you were like, hey, can I show you something? Can I use your likeness from this reel that you posted? I was like, sure, man, go ahead. And you quickly sent me back a few seconds of me talking, and it looked exactly like my reel. And it wasn't my voice. But if I was scrolling really fast past that in a social feed, I would think it was me. And then the most wild part of all of that was you told me you trained it on just a screenshot of the reel. You didn't even use a video. And that brings me to this report you guys put out from captions this week.

5:07And it's really digestible. It's about kind of the state of deep fakes and the technology and where it's going. It's not like your average academic research paper. I encourage everyone to check it out at the Mirage app website, mirage.app. The title is We Build Synthetic Humans. Here's What's Keeping Us Up at Night. And I want to get into that. I like the way you construct this with the four generations of deepfakes and the technology there. So maybe we can start with that. Can you kind of walk us through when deepfakes started, how they worked, and then getting to where they're going. Totally.

5:46Yeah. And, you know, completely, obviously, agree with you on how quickly the technology is developing. It feels like every week, every month, you know, there's like actually miles and miles of progress that's happening, you know, in this space. And it's not the type of thing that you really talk about too much, especially in this context, right? And people have been quiet about it in some ways, right? Obviously, everybody wants to talk about the positive sides. But we wanted to make sure that we're also like, you know, being honest and open about like what we're discovering in terms of like, what are the negative side effects and use cases?

6:18And just so people know, you know, what's out there and how we might address it, right? Both our own company and others in the world as well. That being said, you know, on the deepfake front, like with the report, we really tried to structure it as like not from a technology perspective, more from the user perspective. Like how do people perceive these deepfakes, right? And how they might classify these things, right? To the everyday person, they might look at two or three different videos generated from different models across four years and feel like, hmm, they all look kind of real. I don't understand what the difference is.

6:53But they actually end up being slightly different in the way they're generated. And based on that, there's tells of how you can tell, okay, this is generated versus this is real just by looking at it. And with our own company internally, we look at these videos all the time. Whether it's our own technology or whether it's other companies that are out there, and there's many at this point, right? So we've come to the point where there's many people on our team who can look at a video. And just by looking at it for a few seconds, they can be like, oh, this was generated by this version of the model of this company, probably around this time period, right?

7:27Because you just know that there's certain things off about it or certain things that, you know, are just the telltale signs. But, you know, that's starting to happen. the interesting part of that is like i see that actually across you know all the different generative models these days like i can look at a piece of text made by chat gpt and immediately know that it's chat gpt like within three lines or four lines because it says one of the things that it always says like it's not just x it's y you know like or there's an m dash the m dash is exactly the classic one yeah yeah so i think those same things are happening in the video space as well, where you can kind of tell.

8:03But at the same time, as the technology is developing, the whole goal is kind of to make sure that there's no tell, right? Like that's kind of the point of the technology itself. And you increasingly can't tell right now. That's, I think, the thing that freaks me out is, and you see it almost reflexively, where if something looks iffy now, you look at the TikTok comments, and it's people just saying AI, AI, AI, over and over and over. So like people are starting to see this, but I think it's also getting harder. And that's, it's interesting that that's happening at the same time. So I want to get into the state of things right now, but I guess I think this audience will find this interesting.

8:41These four generations of deep fakes, can you explain kind of the first one and then how that progressed to the fourth one that you all talk about that is coming? We classified it in terms of how much of the video is actually generated, right? And that is actually what has been evolving right now. Now, within each generation, there is development of the technology of that generation that still happens. So, for example, the first generation we see is like face swaps, right? So let's say you want to impersonate Tom Cruise, something that actually happened. If you wanted to do that like four years ago, you would find a Tom Cruise impersonator, basically.

9:15Like a real person who looks and talks like Tom Cruise, says similar things, has a similar voice, acts in a similar way, probably is the same build, height, and weight. And then you would have them basically, you know, act a bunch of stuff. You would record it. And then you would use this face swap technology to basically place Tom Cruise's face on top of that. Now, when this technology first emerged, you actually needed to train on a bunch of Tom Cruise videos. So you probably would like download a bunch of his publicly available videos and train your model specifically to be able to emulate Tom Cruise.

9:47And then you would copy the likeness of him on top of this body double. And that's kind of what you get, right? Right. And that tech got pretty good pretty quickly. There's actually accounts that got pretty viral for doing that. There's a viral TikTok, the deep Tom Cruise that you have in the report that I remember at the time. Yeah. And actually today there's accounts that are impersonating Leonardo DiCaprio and a bunch of other people. These are on Instagram. I mean, if you go find them, they have like some of these videos have 50 million to 100 million views on them, which is crazy. But they're using the same method that the generation one method.

10:22But that's happening today. One of the things that's changed on that, and I think that's evolving, is today's methods of face swap don't need to train on a particular person. They're kind of generalized. So they're using more recent technology that lets them basically just take a picture and face swap somebody without actually having to train specifically on someone. So even if there's no footage available, sometimes that can be the case. There's people who are still abusing this technology today. Now, the problem with this particular generation is that it's actually much harder to tell than other ones.

10:57The reason is that most of the footage is real. And if most of the footage is real, then obviously it's going to look real. What's not real is just the elements, the features of the face. right? Which is a small component, right? The voice is real. The lip movement is real. The expressions are real. The hand gestures, the body, everything is real. What sometimes can be an obvious tell for this one is you didn't get the perfect body double, right? So if you look at the Leo videos in our report, you'll see that the person, you know, from a body perspective doesn't look like Leo. And that's how you can tell that it's obviously not him, right?

11:38But the Tom Cruise one is like pretty spot on and that makes it difficult to be able to tell even a really trained eye may not be able to tell if someone did something like that for a malicious purpose right and oftentimes like these types of models are being used for like you know obviously using the likeness of somebody without their permission i i don't think tom cruise or leo have given permission i imagine maybe i'm wrong right i think definitely not yeah and that that goes into the second generation, which is this lip sync on real footage trend. We saw some of this with even the candidates in the last US election.

12:12And you all say in the report, this is the quote, first time you could make someone that exists say something that was never said. Exactly. Which is a wild statement. And this is relatively recent. Yeah. And this is being actually like, you know, many articles have talked about this. There's been a lot of press about this as well, but like these technologies are being abused. I I mean, go look at any of Elon Musk's Twitter posts, right? And you'll see the first few replies are going to be these fake videos of Elon Musk saying, buy cryptocurrency or something. You know, I'm giving out cryptocurrency, come to my website.

12:46Like something like that, right? All of those, it looks like he's speaking it, but these are actually lip synced with one of these models. And there's actually like 60 to 70 providers of these models today, you know, which is also wild. A lot of them, most of them are not American based. They're like in China or other countries where, you know, there's not really much we can do about how they police it. Well, there's not much we can do in the U.S. either. Let's be clear. That's fair. At least we have some level of control here. Yeah. Maybe we can pass a law or something. But yeah, there's very little we can do with some of the foreign ones.

13:19And a lot of these ones have zero moderation, right? And they have no intent to moderate. No, they don't even have a positive. They want it to be used for fraud and other types of purposes. Like, that's okay for them, right? Because it makes money for them. They don't care. Can we name in shame? What are the worst ones? I think the worst offenders are the Chinese ones because there's no consequences for them doing that, especially abroad. I imagine they have laws locally to prevent that. But, yeah, the American ones, I think, mostly are doing the right things where possible or at least have the right intentions.

13:54So yeah, I would say even the largest Chinese ones, I mean, go look at like Dream Enough, for example, it's a ByteDance thing. They offer lip sync with, you know, zero moderation. Like you can make pretty much anything you want. And there doesn't seem to be any intent to moderate either, which is interesting. So - And you all have a TikTok video in this section of the report that's fake that is very believable because it's the classic recipe thing where it's just a cutout of the person overlaid on top of the recipe. and I would never know that this was fake if I saw this in my feed. So this was actually a lesson for us because people were using our platform to create certain types of videos as well that we ended up finding out about.

14:37We took a bunch of steps to prevent that type of abuse. It wasn't world-ending misinformation or anything like that, but it was things that were misleading. And so we wanted to make sure that we are covering our bases on those types of videos as well. So that's how - I want to get into those. I want to get into that, what you guys did and the challenges you're facing there for sure. Let's go quickly though to this third generation. So this is full frame generation and short form video. And these are the newer diffusion-based models that for the first time you can generate everything from nothing basically in the frame.

15:13Exactly. With no existing footage that you're relying on. And I think for a lot of people, this is VO3. This is Google's latest model that they've seen. You can obviously tell that it's fake. But you all, I think, outline how for the first time, this is what's really kind of potentially scary about this and also cool, but definitely scary is that it's this, you know, generating everything from nothing concept. Exactly. And like, there's no base footage, right? Like essentially, especially from like a press perspective, right? If you want to like prove that something was generated, right? And show that, oh, this was like a deep fake created of like XYZ candidates saying ABC that they didn't say, right?

15:54You can always find the base footage and be like, here's the original. So clearly, you know, this was manipulated, right? But with these sort of transformer models now, which are generating the entire frame, right? They're not, there's not, there's no original footage to rely on, right? So there's nothing at the base of it. there's no thing you can find to prove that oh look at that like this is the original which is what makes these you know more more difficult to detect in the long run you know some of these technologies are still developing and so the fidelity may not always be there and all the models that kind of work in this domain but that's just a matter of time right more data and more training that will get really good and once that's there once the fidelity is there right then you kind of end up in the situation where it's really hard to tell and you can not only have people say stuff they didn't say you can put them in locations they weren't in or doing stuff they never did while saying things they never said right so it's a whole other level of complexity from like a detection perspective from a moderation perspective and you know obviously that applies to celebrities and public figures and some of those more abused personas i would say right but it also applies to like everyday people right you don't want like someone to take your friend and put them in some random place and have them say something either right like that's not that great either we need to take a quick break we'll be right back

17:24support for this show comes from linkedin imagine if any of the movies that included the line I need the right person for the job, settled for I'll just take about anyone. How many heists would have failed? How many deals would have fallen through? How many secret spy missions would have ended in disaster? So why would you accept just anyone when hiring for your business? When you need the right person for the job, you can turn to LinkedIn Jobs. And now LinkedIn Jobs is stepping things up with their new AI assistant. So you can feel confident you're finding top talent that you can't find anywhere else.

17:58With LinkedIn Jobs AI Assistant, you can skip the confusing steps and recruiting jargon. It filters through applicants based on criteria you've set for your role and surfaces only the best matches so you're not stuck sorting through a mountain of resumes. Hire right the first time. Post your job for free at linkedin.com slash partner. Then promote it to use LinkedIn Jobs' new AI Assistant, making it easier and faster to find top candidates. That's linkedin.com slash partner to post your job for free. Terms and conditions apply. Support for this show comes from Odoo. Running a business is hard enough, so why make it harder with a dozen different apps that don't talk to each other?

18:47Introducing Odoo. It's the only business software you'll ever need. It's an all-in-one, fully integrated platform that makes your work easier. CRM, accounting, inventory, e-commerce, and more. And the best part? Odoo replaces multiple expensive platforms for a fraction of the cost. That's why over thousands of businesses have made the switch. So why not you? Try Odoo for free at odoo.com. That's O-D-O-O dot com. Having a smart home is a cool idea, but kind of a daunting prospect. You have to figure out which devices to buy, how to connect them all together. It's all just a lot. But for two weeks on The Vergecast, we're trying to simplify all of it.

19:32We're going to spend some time answering all of your questions about the smart home. And then we're going to go room by room through a real house, my real house, and try to figure out how to make it smart and how to make all of that smart make sense. All of that and much more on The Vergecast, wherever you get podcasts. This special series is presented by The Home Depot.

19:55We're back with CaptionCO Gaurav Misra. The tell for this right now is basically just that it, besides the fact that you can clearly tell like with VO3 that this is, it looks crazy. It looks fantastical. Usually it's AI, but there's this eight second mark limit, right? Because that's the model output limit because of, I'm assuming, compute constraints. But that's like for VO3, for example, you can't do more than eight seconds. Right. Right. So this is all about the window that the model is trained on. So these models, the diffusion model specifically, like they're trained on a specific length, or at least this generation that exists right now.

20:31So that means that they're just trained to produce eight second videos. They can't produce nine second, seven second, nothing else, right? So it doesn't have to do with compute constraints. It's actually just the constraints of how they were trained. So there is compute constraints. So essentially as that window increases, so if it's nine seconds, 10 seconds, 12 seconds, there's a quadratic or like a square increase in the amount of compute needed. So eight seconds versus 16 seconds is going to be four times the compute to run it, also to train it probably. And then same if like you double that, you quadruple the amount of compute, right?

21:03So that's why all of them kind of turn out about this eight to 10 seconds, because it just like scales too fast after that. And there's not GPUs that can like, you know, run that quickly enough for it to be viable. So those things are being solved right now. I mean, we're working on, You know, obviously a next generation model that is not going to have any of these constraints and we can bypass basically all the compute constraints around this and create pretty much unlimited length video with like no compute restrictions on length and stuff like that. I'm sure other people are working on it too.

21:34And so this is kind of where the fourth generation will start, right? And with the fourth generation, we think there's properties that make models the fourth generation. None of the existing models actually capture all of those properties, but there are models today that capture parts and individual elements of what a fourth generation model would look like for this. So for example, unlimited length is a critical one. And we actually just released unlimited length a couple of days ago. It's not publicly announced right now, but it's available to certain users in the app. And so you won't have like eight second or 10 second limits or anything.

22:11It can do continuous. I mean, the social post that I shared with you was a continuous shot, right? There was no cuts in the middle of it. Once I was speaking, I was speaking throughout the video. I think it was 50 seconds or so. And I've posted a couple more of those types of videos recently where you'll see there's no cuts. It just like continuously speaks through the video. And even in the deepfake report, when we share some of the fourth generation results, you'll see that they're continuous videos with no cuts in between. So they can go to arbitrary lengths, essentially. And then you'll see models like VO3 excel at another property that we think is going to be important for these four generation models, which is multi-person interaction, because that kind of is the next frontier of a lot of this type of stuff.

22:50I would say on top of that, another property that we would probably see is not just character consistency, which, you know, is already being achieved pretty well. It's pretty close, but also location consistency. So let's say there's two, three shots, different camera angles, looking at the same person, the same location, the location remains consistent across those shots, right? So if there's a glass on the table, when the camera cuts to the left, the glass still is on the table at the same location, right? And it's the same glass, right? Those types of things are all being worked on. Like these things we're working on, I'm sure other people are working on as well.

23:24You can imagine the positive value of something like that. Like you can literally create completely new environments. You can literally like shoot a movie if you wanted to without even moving a single item without even leaving the room. You can come up with a location, tables, chairs, whatever you want in there, the people that are sitting in there, what they're saying, all of it just created out of nothing. Just like in civil engineering, you can imagine a bridge, you can't build it. In software engineering, whatever you can imagine, you can build it. Essentially, movies and TV and video becomes like software engineering.

24:00Anything you can imagine, you can just do. There's nothing physical to stop you from getting it done, right? And so that is the power of what's going to get unlocked in sort of this next generation of models. Obviously, there comes to the downsides. Yeah, and you all have this example where you generated Jon Stewart reviewing Dune 2 in a military uniform. And the audio is not right. And the audio, I think, remains a huge tell, even with these frontier synthetic models. But if the audio was off and it was just the captions, and I scrolled past this, I would definitely think this was Jon Stewart.

Read the full transcript

24:35Totally. That, I think, is the biggest tell of the fourth generation borderline models that exist today is the audio, right? Because audio models are obviously making progress of their own, but none of them yet focus enough on sort of like casual, everyday voice that sounds like it was recorded very ad hoc. Like a lot of them come off a little bit too professional, a little bit too advertising or podcast or something like that. These models struggle with that because they were just never built for this type of, you know, casual use case. And then secondarily, you know, there's an issue for matching audio and video, essentially, where let's say you're on the streets of New York City in the video, you're walking down, you expect there to be some sort of noises is happening.

25:27You know, you expect there to be like people talking and walking in the background, maybe cars honking and driving by and things like that. And so audio models just don't have that context today, right? Like when you're generating a piece of audio, you just can't tell it like, actually, this is a New York city on the street. So make it right for that. Right. It also carries through into things like, you know, if you're walking down the street of New York city, like you're probably don't not having like a professional microphone in front of you, right? So audio is not going to sound perfect. It's going to be like, kind of like sound like it's recorded at a distance.

25:56And so those properties of the audio can't be generated right now. And so because of that, it's very obvious. Sometimes you listen to the clip and it looks like it's recording a podcast studio, but you watch it and it's walking down the streets in New York City and your brain is just like, this doesn't connect, right? And so that's like the key problem. I would say Google's VO3, from what we know about it, what people know about it, they haven't really released any details about this, but supposedly they're generating audio and video simultaneously in a single model, which is definitely a technical achievement if true, because people haven't done that before.

26:33And if that is true, then it is doing the right thing. The audio is being generated with context of what the video is going to be. That's why it's able to generate background noises and things like that. It just feels to me a little bit under-trained right now because it has a little bit of the, you can hear sort of cracking in the, in the audio. As if you look, listen to it closely, you can hear like, you know, almost cracks between the voice. You can, you know, hear a creak a little bit. And that's like a tell I would say right now, but I imagine if it's under training, I'm sure they'll figure it out in like a couple more months.

27:06So yeah, I was going to ask, when do you think the audio realness catches up to the visual realness? There's two paths. I think we can look for, to the audio companies to figure that out. And there's companies like Eleven Labs that are working on, they have a V3 model now that just launched, which was maybe a month ago. They previewed it. It's more widely available now. That seems to be much closer to a realistic audio than anything we've seen. No other audio company comes close to that, I would say. But it seems like they're still working on making improvements on that to make it even better.

27:38On the Audio Plus Video side, Google is the only company that's working on Audio Plus Video at the moment that's publicly released. We're also working on it internally. But we haven't released anything about it and probably not for a little bit more until it gets pretty good. So that's kind of the state of the world as we know it right now. But we think that over the next six months, this focus is going to come back because it's just use case driven, right? Like all these technologies at the end of the day, like, you know, there's easy solutions to, you know, we talk about like, okay, there's deep fake audio, there's deep fake video.

28:11Oh, it can cause damage. What's the easiest solution? Let's all stop working on it. let's all stop working on it. That's one solution, right? It's the most extreme solution. But at the end of the day, there's commercial value in this, right? And just by the nature of capitalism and how much commercial value there is and all this stuff, it will happen, you know, and nobody can stop it. No law can stop it because there's other countries, you know, the world is a big place. And when there's so much commercial value in this type of stuff, things will move. And so with audio, it just comes down to like, is there commercial value in having a conversational sort of casual style, social media style audio generator?

28:51Right. I think there is, I think it's going to be realized, you know, in the short term at some point and someone will do it. I feel like the big question overhang all this, which is Grav, like you're building this. You're also talking about the risks of it and how it's already been used to deceive at scale. Why are you building this? I know there's commercial interest. I know that, you know, you've raised a lot of money, like you're making money. I know that there's a lot of interest in this technology for a lot of things, but God, when I look through how this is being used already and how I feel, even as a relatively, I would say, informed internet user, which is that I am increasingly unable to quickly decipher what is real or not, But why is this worth it?

29:38Why is it worth building this when you're seeing the impacts of it already? It's a great question. So, I mean, I think I fundamentally believe that the technology itself is inevitable, right? Like, we cannot escape the development of this technology. It's kind of like trying to stop LLMs from happening, right? They will happen. We may or may not be a part of it. We can choose not to be a part of it. But the technology itself is, it's already on its way. and just by the nature of capitalism, it will get built, it will get adopted, right? If the value is there and it seems like the value is there and that's what we're seeing.

30:12Like we have millions of users using this stuff to literally do, you know, stuff like getting their business off the ground, right? Like it's a one person shop, like a lot of our users are like, by the way, millions of people who are small businesses like nail salons and like, you know, oh, it's like a plumber or an electrician or it's like a home cleaning service or, you know, trash removal, lawn services you know these types of users right like everyday people that are using our technology to help get leads sell stuff make a name for themselves grow their business things that are just creating a ton of value in the economy so it's hard to argue with this type of value i think right on the flip side you get these negative effects and the way i think about that is like i mean so many times in today's world like you look at the state of the world and you imagine I'm like, I wish things were different.

31:04I wish there was something I could do about some war in some country or like some political thing that's happening. I wish I could contribute to it, but all I can do is I can vote or I can donate money. And that's really the extent of anything that I can actually truly achieve, right? As an individual. And so when I get a chance to be actually at the front and center of helping decide where a certain inevitable technology is going to go, how society might change and actually have an input into it, I'm going to take a seat at the table, right? Like I'm not going to walk away and be like, you know, not for me.

31:38I'm going to take a seat at the table and be like, I would love to help figure this out, right? I would love to work with all the information I can get, work as hard as I can to make sure that this is done in the right way and that the positive value outweighs any negative value in the long run, right? Hopefully. And by the way, that's the case with every technology, right? And we've seen this over and over again, how there's positive and negative. There's almost every technology that's been developed, right? The question just comes down to, do you want to be part of the conversation or do you want it to be figured out by other people?

32:10The title of your defake blog post this week is what's keeping us up at night. Is there a way that captions could be used that if you woke up and found out would make you reevaluate all this when you're weighing the good and the bad? I would hope not because we are staying up all night to make sure that that does not happen basically, right? And we are constantly thinking about strategies, constantly thinking about what a future world might look like. Fast forward a hundred years, right? Like what does the world look like when all this stuff has been fully developed and played out. Backtrack from there, what went wrong?

32:48You know, what are the things along the way that were learned? And how can we make sure that mistakes are not made and that things are done the right way? You've done that exercise, like thought out that far? Yeah, definitely. Lay it out. What is the worst scenario? I would say one possible scenario, right, which seems most likely given the current information that we have, right? Like, we don't know what technology might develop around, you know, identifying these videos or so on and so forth, right? So given a lack of that information, let's say nothing changes in any of those areas. I think the most likely scenario is that we kind of go back to how things were 100 years ago.

33:32Essentially, video didn't exist. Audio recordings didn't really exist. Photos barely existed. and so let's just call it they didn't exist everything was practically just a matter of trust right like what you believed and what you saw just depended on oh well who wrote it well it's this publication the xyz publication and they're very reputable you know it's like oh okay so i believe it right or maybe it's like a trial and it's like well did they do it well we have three witnesses. Do we believe the witnesses? Well, yeah, they're upstanding people, you know, so we believe them. That's sort of the worst case of where things end up, right?

34:14You mean like society wholesale rejects this and goes back in time? Essentially, because nothing can be trusted, right? Like any kind of information digital, right? Like whether it's audio, video, written, right? Like nothing can be actually recorded. So the only thing that can be recorded is what people have seen, right? And that's the only that can be trusted, but that can be trusted only as much as you trust the person, right? And everybody has motives to manipulate the truth. So that's obviously something that's the way, but that is how the world worked a hundred years ago, right? That's exactly how the world worked.

34:44Not to say that that was a better world. I don't think so. I think we've come a long way from that, but that is the worst we could regress to, right? So that's like the base case, I would say, right? Is that world. I would say the one difference between that world and today is that information can travel much faster today than it could back then, which means that damaging information or information that, you know, could potentially have any harm can make its way through a lot more people today and probably in the future, even faster than it could a hundred years ago, right? The hundred years ago, it might've been limited, but at the same time, there's a flip side of that too.

35:22A hundred years ago, if something made it out, if, you know, some damaging information, information made it out across the U S like, oh, something about a candidate or an election or something, it will be really hard to counter that because you don't have a way of reaching all those people that quickly anymore either, right? Which is flipped now where you can counter that and be like, hey, this is incorrect information, right? Like this is the reality. And so, you know, it kind of falls in a balance of pros and cons between, you know, where things are today and where things were back then. But I would say that's like the worst case scenario.

35:56And then building a topic - case scenario is not just like society collapses because we go into echo chambers and we don't trust each other and we go to war with each other that's that seems reasonable fair yeah i believe in people in society enough to think that if this scenario has played out in the past and it has literally right like there was a time where you couldn't trust anything but what people said then it will probably play out in a similar way in the future if that is the reality we come back to where, well, you can't trust anything but what's been said by people and how much you trust them, right?

36:32Which actually is interesting because it brings back almost an importance of the press and trusted people and journalism and all the things that have been reversed. I mean, pendulum swings, right? So I could easily see the pendulum swing the opposite way 20, 30 years from now, where there's a world in which this is all solved by just having trusted journalists and, you know, trusted names and, you know, that type of thing, some sort of system, you know, around that. And the pendulum swings the exact opposite way of where it is right now, because right now it's like nobody's trusted basically. And so it's very scary to me that trust in institutions are at an all time low as deep fakes are proliferating.

37:13It's true. And I think you're saying that people may go back because they realize like, oh, wait, when I don't trust anything and nothing is actually real and I can't tell, I actually just have no way to make sense of the world. Yeah, I think that's one possible outcome. And I think, you know, with video and audio and sort of these levels of deep fakes appearing today in a world where there isn't already that much trust in anything, it might accelerate that trajectory to going back to the world that existed before, right? A pendulum swing in the opposite direction. That's one scenario, I think.

37:53But I think that's the base case. And that's the case where we can't figure out anything. There's just no solution that we can come up with. I would say there are clear directions of where solutions could be heading. There's ideas like crypto signing and things like that. I'm not talking about the blockchain, but like just signing cryptographically videos that are recorded on device as they're recorded with the identity of the person who recorded it. so you know where it was recorded when it was recorded who recorded it as opposed to like an unsigned video which is like it could have come from anywhere there's no source there's no crypto signed sort of like signature to say that this was recorded by any particular person right so you might have some concept of like trusted videos right which were recorded on crypto signing devices right where maybe on an iphone there's an inbuilt like chip that automatically is signing the video as it's being recorded with your, you know, your name or something like that.

38:48So, you know, who recorded it and location where it was recorded. So things like that could emerge as, you know, a possible outcome in the future. Isn't that just fancier metadata that can be scraped just like metadata can today? Well, it is metadata, but it's cryptographically signed. So it's one step further. I think the problem that I have with today's metadata, like SynthID and like some of those other solutions. Yeah, you're not a fan of watermarks. says in the support, you don't think that's going to work. I think it's just a very thinly veiled attempt at like avoiding the problems. Like it's like the answer you want to give to a journalist when you're like, well, what are you doing?

39:24Oh, we put a watermark on it. It's like, yeah, it's so easy to remove the watermark. It takes like five seconds. And by the way, it's easy to accidentally remove the watermark. You know, you may not even be wanting to remove it. You accidentally, you crop the video and it removes the watermark. Right. And so I don't know, what's the point, right it doesn't it doesn't seem to be actually doing anything besides just like pretending that it's solving a problem and so i'm not really a big fan of that and you know i'd be curious to see like i can't come up with how how exactly they would make it better than what it is right now maybe they have a roadmap and i'd be curious to see like the companies that are working on that kind of stuff you know where where something like that might go but it's not to my satisfaction i would say i will say that maybe there are cryptographic solutions though right like that that that might be an interesting direction.

40:10It's still watermarking, but it's like crypto watermarking. Maybe there's something to that. There's obviously like things there too that could go wrong in terms of like on device, people could hack stuff and all kinds of things. But it seems much more reasonable to me if it was crypto signed rather than just like, oh, I just threw some metadata on there that anybody can change literally, right? Or remove, or I could myself literally remove proving that it's false, right? I would also say a big part of it is also like, do you want to mark videos that are real or do you want to mark videos that are fake?

40:47You know, which one is more important? I haven't thought of that. Yeah. I kind of feel like you want to mark videos that are real because then at least, you know, when one comes along and you got to mark them in a way that it can't be removed, obviously right or or manipulated but at least in that world it becomes clear that maybe the social media platforms or something can have some something that says like like a checkmark right like a blue checkmark but you know that means nothing anymore but if that meant something like a blue checkmark on the video that's like this is a real video it's verified through crypto right as opposed to like this is an ai generated video which i don't know if that really adds any value.

41:28We need to take another quick break. We'll be right back.

41:39Support for the show comes from Rippling. If you're a business owner, here's the truth. SaaS promised to make work easier, but now the average company is buried by hundreds of apps that silo your teams, slow you down, and simply don't work together. That's not SaaS. That's sad. Software as a disservice. That's why you need Rippling. Rippling is the unified platform for global HR, payroll, IT, and finance. They've helped millions replace their mess of cobbled together tools with one system designed to give leaders clarity, speed, and control. By uniting your employees, teams, and departments in one system, Rippling removes the bottlenecks, busywork, and silos your software created.

42:23Automated, perfectly in sync, and seriously simple to use, Rippling gives your company one source of truth for your people, their data, and everything they touch. With Rippling, you can run your entire HR, IT, and finance operations as one. Or pick and choose the products that best fill the gaps in your software stack. And right now, you can get six months free when you go to rippling.com slash decoder. Learn more at rippling.com slash decoder. That's rippling.com slash decoder for six months free. Terms and conditions apply.

43:02Support for this show comes from Rocket Money. There are very few easy things about managing your money. How many subscriptions is too many subscriptions? How do I create an easy and effective budget? How do I keep track of all my accounts? Should I just hoard all my cash and bury it in the backyard? With so many ways to mismanage your money, sometimes it feels like the easiest thing to do is just ignore it. But that's where Rocket Money comes in. It tracks your spending, your accounts, and your subscriptions all in one place and makes managing your money actually easier than pretending it doesn't exist.

43:37Try Rocket Money for free at rocketmoney.com slash vox. When it comes to sales, efficiency is the name of the game. Every minute you spend manually entering data and chasing down documents is a lost opportunity to close the deal. That's where PipeDrive comes in. PipeDrive lets you supercharge every sale by automating the process for your entire team. From scheduling sales calls to email marketing, it's a powerful, simple CRM built by salespeople for salespeople. Join the over 100 ,000 companies already using PipeDrive. No credit card or payment needed. Just head to Pipedrive.com slash Vox to get started.

44:19That's Pipedrive.com slash Vox and get started with a 30-day free trial.

44:29We're back.

44:33So that sounds still kind of theoretical. In the absence of the cryptographically signing something on device, What are you doing at captions as someone who is literally building these models and releasing them and having millions of creators use them, as you just explained? What are you doing to keep your models from being used to do bad? yeah i mean so obviously we've run into these problems before like we talk about in the report how you know people have tried to abuse our models before and it's not a thing that happens like you know every once in a while there's often people who are trying to abuse our models all the time right and the thing is though in today's world with a few different options available on the market and some of the options being completely devoid of any kind of moderation or any attempt to even moderate, if you make it just hard enough, people will abandon it and just go to the place where it's absolutely easy to do it.

45:33Right. So that's kind of why we haven't dealt with this problem as publicly and as obviously as other companies might have, because we make it just hard enough where they just choose to leave and go to the platform where it's really easy. That's kind of what we've seen more recently, more specifically on, on what we are doing at this moment with some of our most recent models that are actually mentioned in the report. We're starting off with like, uh, the most basic approach, which is let's just, let's just look at all the videos created. Let's just have human look at all the videos created. And if something You can do that?

46:13You're at the scale where you can have humans look at every video? At this moment, yeah. But that may not reasonably be possible at some point in the future. But at this moment, even though it might be time-consuming and operationally complex, that's where we want to start. Just so we can understand the patterns, right? We might not even understand all the types of abuse that are possible, right? There's things that people come up with that we wouldn't even have thought about. One of the ones that I've been thinking about recently is like ID verification abuse, right? Where you kind of take a person's likeness, have them hold up an ID that's like, you know, obviously digitally created one of themselves and use that video and have them say something like, this is me on this date or something like that.

46:58Use that video to get verified on, you know, some sort of verification system. Many companies have these, right? Like clear type stuff, basically. and yeah i mean these are the types of things that every day people will come up with new ways to abuse right and it's it's an ongoing and constant battle right it's like we do things we wait and watch we see five new things appear and then we do more things and then we wait and watch and then five new things appear and people figure out how to bypass stuff people figure out how to like you know get squeezed past the defenses basically right and then we have to continue improving, basically.

47:34And that's like never ending. You know, there's no end to that today or ever, really. But I think this is a pretty played out playbook in that sense, right? Like, this is kind of how trust and safety teams have always operated. And this is kind of how abuse has always been prevented. It's like always an ongoing battle, right? Like, there is criminals out there who will break the law, you know, and we have to do our best to make sure that doesn't happen. Yeah, but as we've talked about, this is also very new, like the ability to deep fake someone else to impersonate someone else to commit fraud in ways that weren't possible before and like for example can i can i go into mirage studio right now and deep fake anyone and is what is the constraint on what i can make them say what i can make them do do you even know that that's me or it's not me like that's that's what i'm trying to get at Totally.

48:25Yeah. I mean, there's like unsolved problems around like, how do you make sure that someone can't upload a picture of somebody who is not a public figure, who they might know in real life, but is not a public figure. So we wouldn't be able to identify them and then have them be able to say something, you know, that they didn't say. And in those cases, we have to depend on like people complaining. And that has happened where people reach out to us and said, hey, like someone made a video. I never said these things. And we have to go figure out who did it, ban the accounts, reach out to the, you know, share the.

49:00Usually there's a lawsuit involved between the parties. So we're obviously happy to share whatever information is necessary for that. We have to take those on a case by case basis, basically. Right. But it always comes out as like a user complaint or, you know, something of the nature. Yeah. I mean, in your video, your deep fake video accompanying this report, you say, you know, please use this tech responsibly and quote, keep us honest. And to me, it's kind of like handing like teenagers keys to a Ferrari without having, you know, an ID yet and saying like, keep us honest, like go have a jewel ride, but keep us honest.

49:35There's no forcing function on you or the industry at large to actually have responsibility aside from like, yeah, you want to stay in business. But as you said, there's a lot of market pressures with people doing way worse things and allowing people to do way worse things with these models. So you're already in this, like, as you said, hyper-capitalistic setup. And there is zero regulation on this topic, really. I mean, there is around deepfake nudity, which I was really glad to see. but what do you think needs to happen beyond that to actually give you you and your industry actually a little more liability in the outcomes of this technology and not just be like here's the keys you know don't crash i think the first step in my mind and i'm always open open to hearing feedback on this from other people and you know invite viewers to reach out to me too but i would say the first step in my mind is to have laws around criminality of using people's likeness to say cause harm basically because there's not criminal laws around that right now there's civil laws around it you know i think a well-crafted law around some sort of criminal action around using people's likeness against their will for certain things i think that might be the place to because that doesn't exist right now.

50:57There is a civil liability, but doing lawsuits and stuff is just an expensive thing. And it may not be accessible to everybody who's affected by something like this. So I would say that's the place to start. What is the number one commercial use case that you're finding that people are using your models for right now? So I would say the number one commercial use case is marketing. Okay. So that's when we're talking about like the benefits outweigh the risks. Right now, people are using this to jumpstart their marketing efforts and do things with marketing they couldn't do before. That's right now you would point to as the top benefit of this technology.

51:39The top benefit today of the technology is being able to make videos for marketing, for sales, for entertainment, whatever that might be. Essentially, that means that people can kickstart their business out of nothing, right? Like people can make money, make a living, you know, feed their families and do stuff they couldn't have done before, right? Which moves, you know, not just their own lives, but it moves the entire country and the whole economy, right? So I would say that that is the flip side. And creators are making AI versions of themselves with Mirage. You had that release a few months ago that went pretty viral where you were showing all these examples of that.

52:22And that's, I mean, they can be in the ads, but when you think of traditional ads, it's not necessarily like generating the person in the ad. But is that really the main way people are doing this is they're putting AI versions of themselves in the ads? Yeah. So what has been popular is putting versions of yourself. So a lot of times, you know, think about a lot of these businesses that I kind of mentioned to you, right? they're not exactly camera ready, right? Like these are not people that are, oh, let me just like show up in front of the camera. Like I'm like, you know, some famous Hollywood actor or something, right?

52:57They're doing this for the first time and it's like kind of daunting. Also, it just doesn't come naturally to everyone. It doesn't need to. And that often holds them back from like having their business take off, basically. So think about this, right? 20 years ago, everybody needed a website. And there was an age where literally you could not have a business. If you had a business, you needed to have a website, right? And then after that, 10 years ago, maybe was the age of design. Oh, it's not enough to have a website. You have to have a design, right? Without a design, great design, like no one's going to trust your nineties looking website.

53:31And then this is the age of video, right? Like literally everything sells through video. Everything markets through video. TikTok is obviously a big platform. Reels is huge now. YouTube Shorts is catching up, right? These are massively growing platforms where the entire economy is running through these things, right? And all the discovery of products and services is happening through these platforms. If you can't participate in this, you are getting left behind. And these types of tools let everybody participate. They let everybody participate in this economy that is a massive economy, where whether you're a plumber or electrician or whether you're like a store or, you know, lawn service or whatever, you need to be a part of this economy to be able to sell, to be able to get leads, to be able to convert people, you know, grow your business, make money.

54:21This is a part of the game now. These tools allow people to participate in that without having to have the skill of being some like perfect person on camera, right? I think that's the game changer. People love what we offer because of that, right? And we have one of the largest number of subscribers of any AI video company because of that. So that's what I would say, you know, has, has changed over the years. Like video is kind of like you live or die on it practically. With AI, I'm generally not as doomerism as a lot of people in the press. I think a lot of this stuff will be sorted out. It will be very chaotic, the job loss, all that stuff.

54:59But that stuff to me will sort itself out over time. This deepfake stuff really worries me. And we've talked about all of, you know, where the tech is headed, how you could potentially fight it. To be honest, it seems like there really isn't a plan. It seems like at scale, we still have to figure this out. So can you leave us with a sense of, you know, as this is progressing in the ways we've talked about, where in the next six to nine months, there's going to be, you know, potentially unlimited synthetic video with audio that matches. How do we navigate this? How do we discern what is real or not?

55:35So I would say the number one thing to do in the short term and the fastest way to kind of mitigate this problem, at least, is to create awareness around it, right? Because I think a lot of people just don't know. Matter of fact, like a lot of that is already happening through platforms like TikTok. I mean, you said it yourself, like anytime there's something off in the video, there's 10 comments there like AI, AI, AI. And that's a small number of people who can tell. A vast majority probably can't, but there's some people who can. and TikTok essentially is giving them, you know, likes for being able to do that.

56:05You know, that's just the nature again of how TikTok works. But that is creating awareness. And like, you know, you'll also see in those videos, probably a lot of comments being like, oh my God, I could never have told, like, this is crazy, right? I don't even believe it. And now they believe it and now they know what to look out for, right? Because if they didn't even know it was possible, that leaves them kind of defenseless with this type of technology. So I would say step one is awareness, which is why we're all we're talking about it and we're open about like obviously the problems you know both for us for other companies we're also open about you know solutions that we're exploring and where things might go we're also talking about like what's coming next and not just like how amazing the technology is going to be because that it will right but also what the downsides are going to be of these next generation of models too right so we can be ready and prepared with solutions for what is to come right even before and by the way there's companies exclusively working on identifying this type of stuff, on finding solutions to detective fakes and things like that.

57:04We also want to give those companies a heads up of like, hey, by the way, these are the things that are coming and we should figure out how to detect these XYZ new things that are going to be happening. So that's the short-term plan, I would say. In the long term, a solution is needed. I think one nice step would be to formalize the legal process around this. right like i think we should have laws that kind of address this area it's a new area i get it why these things move slowly but i think it will get addressed i think there's motion around that in the government but i think that's just part of the puzzle because it's a global economy right like just having laws in the u.s will not do anything especially if the laws are like hindering progress and things like that right so they have to be done correctly but i would say beyond that there's like a more global solution needed and i actually think there needs to be more of a focus on marking and detecting videos that are known to be real and known to be sourced, known to be correct, right?

58:01Rather than trying to like mark every fake video, right? Because I mean, you can literally make fake videos with CGI, right? Like you can, you don't even need AI for it. You can just use CGI to make fake videos. They look super real. You know, you can put them in Hollywood movies. You can make anyone have, you know, people have been brought back to life in movies and stuff before. So this is not something that can't be done. If someone, a bad actor, a government, or, you know, wanted to do it today, they could do it with no AI at all, right? And so we should be prepared for all scenarios, including ones that are non-AI on that front, right?

58:33But I think the right solution is focusing on the videos that are real, marking them and actually surfacing them on the platforms that are distributing, for example, like TikTok and so on. So that's my short and long-term answer. All right, Gaurav, well, it was good to have this real podcast with you between two human beings. Hopefully we can check back in later and check in on the progress. And as you said, keep you all in the industry building this stuff responsible and honest. So appreciate your time. Yeah, of course. Yeah, great. Great to chat. Thank you.

59:08I'd like to thank Gaurav Mishra for taking the time to speak with me and thank you for tuning in. I hope you enjoyed it. If you'd like to let us know what you thought about the show or what else you'd like us to cover, drop us a line. You can email us at decoder at the verge dot com. We really do read every email. We also have a TikTok and an Instagram. Check those out at DecoderPod. If you like Decoder, please share it with your friends and subscribe wherever you get your podcasts. And if you haven't already, please don't forget to subscribe to The Verge, which gets you access to my newsletter, command line, and a bunch of other great stuff.

59:39Decoder is a production of The Verge and is part of the Vox Media Podcast Network. Our producers are Kate Cox and Nick Statt. Our editor is Ursa Wright. The Decoder music is by Breakmaster Cylinder. See you next time.

59:55Support for this show comes from Odoo. Running a business is hard enough, so why make it harder with a dozen different apps that don't talk to each other? Introducing Odoo. It's the only business software you'll ever need. It's an all-in-one, fully integrated platform that makes your work easier. CRM, accounting, inventory, e-commerce, and more. And the best part? Odoo replaces multiple expensive platforms for a fraction of the cost. That's why over thousands of businesses have made the switch. So why not you? Try Odoo for free at odoo.com. That's O-D-O-O dot com.

1:00:36Support for this show comes from The Home Depot. This holiday season, take advantage of savings on the wide selection of top smart home security products at The Home Depot. The Home Depot has everything you need to make your home smarter with the latest technology and products that let you control and automate your home. And with brands you trust like Ring, Blink, Google and more, available in-store and online, often available with same-day or next-day shipping. So you can protect your peace of mind, whether you're away or at home this season. The Home Depot. Smart homes start here.

1:01:13ever feel like your work tools are working against you too many apps endless emails and scattered chats can slow everything down zoom brings it all together meetings chat docs and ai companion seamlessly on one platform with everything connected your workday flows collaboration feels easier and progress actually happens take back your workday at zoom.com slash podcast and zoom ahead.

From the publisher

This is Alex Heath, your Thursday episode guest host. Today I'm talking with Gaurav Misra, the CEO of Captions. You may not have heard of Captions yet, but by now, you’ve probably seen a video that was generated using its AI models. The company’s Mirage Studio platform lets anyone generate AI versions of real people, and the results are alarmingly realistic. 

Captions just put out a blog post titled, “We Build Synthetic Humans. Here’s What’s Keeping Us Up at Night.” It’s a good overview of the state of deepfakes and where they’re headed. So Gauraav and I sat down to discuss the trajectory of deepfake technology and what might be done to prevent it from being misused. 

Links: 

We build synthetic humans. Here’s what’s keeping us up at night | Captions

Google’s Veo 3 AI video generator is a slop monger’s dream | Verge

Gemini AI can now turn photos into videos | Verge

Trump just unveiled his plan to put AI in everything | Verge

Racist videos made with AI are going viral on TikTok | Verge

Microsoft wants Congress to outlaw AI-generated deepfake fraud | Verge

YouTube is supporting the ‘No Fakes Act’ targeting unauthorized AI replicas | Verge

This Tom Cruise impersonator is using deepfake tech to impressive ends | Verge

Credits:

Decoder is a production of The Verge and part of the Vox Media Podcast Network.

Our producers are Kate Cox and Nick Statt. Our editor is Ursa Wright. 

The Decoder music is by Breakmaster Cylinder.
Learn more about your ad choices. Visit podcastchoices.com/adchoices

More from Decoder with Nilay Patel

All 152 episodes
We are not ready for better deepfakesDecoder with Nilay Patel · 57 min
Listen in VO