Twelve Labs Secures $27M Funding for Advancing AI in Video

8 Apr 2024 · 9 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI Today Podcast: Episode Summary

Episode Title

Twelve Labs Secures $27M Funding for Advancing AI in Video

Podcast Description "AI Today" explores the latest advancements in artificial intelligence, covering breakthroughs and ethical considerations impacting technology. The show aims to demystify AI, making complex topics accessible and engaging for all listeners.

Episode Overview In this episode, the discussion revolves around Twelve Labs' recent funding announcement of $27 million and how this capital will enhance AI applications in the video domain. The focus is on the potential impact of this funding on various industries and the future of generative AI in video.

Key Highlights

Twelve Labs Overview

  • Company Focus: Twelve Labs is a San Francisco-based startup pioneering video language alignment and multi-modal video understanding.
  • CEO's Vision: Jay Lee, co-founder and CEO, envisions tools that allow developers to create applications capable of comprehending videos like humans do, linking natural language to video content.

Technology and Applications

  • Semantic Search: The initial focus is on semantic search, likened to "control F for videos."
  • Capabilities: The technology allows for:
  • Identification of actions, objects, and sounds in videos.
  • Automatic categorization of scenes and extraction of key topics.
  • Creation of video summaries and organization of content into chapters.

Importance of Generative AI in Video

  • Future of Video Creation: The podcast discusses the potential of generative AI in video creation, emphasizing the need for advancements in:
  • Natural language processing.
  • Image generation.
  • Audio integration.
  • Challenges Ahead: The complexity of generating coherent video scenes with appropriate audio and visuals is highlighted, with predictions that significant improvements will be made in the next 1-2 years.

Competitive Landscape

  • Comparison with Other Technologies:
  • Twelve Labs aims to differentiate itself from large tech companies like Google, which use models like MUM for video suggestions.
  • Their model is designed to provide deeper and more customizable video analysis compared to existing technologies.

New Developments

  • Pegasus 1 Model: Introduction of a multimodal model capable of performing various video analysis tasks.
  • Investment and Growth: The recent funding round has increased the total capital raised to $27 million, attracting prominent investors like NVIDIA, Intel, and Samsung.

Applications in Various Sectors

  • Diverse Clientele: Twelve Labs has attracted a user base of 17,000 developers across sectors such as:
  • Sports (e.g., NFL)
  • Media
  • E-learning
  • Security

Future Outlook

  • Jay Lee expresses optimism regarding the future of video understanding and AI, anticipating that Twelve Labs will enable clients to achieve remarkable technological feats in video analysis.
  • The potential for generative AI in video is framed as the next frontier that will be transformed by advancements in technology.

Conclusion The episode emphasizes the significance of Twelve Labs' contributions to the field of video AI and the expected growth and innovation stemming from their recent funding. The discussion highlights the integration of various AI elements and the potential applications that can revolutionize how videos are analyzed and created.

Additional Resources

  • Join the AI Box Waitlist: [AI Box](https://AIBox.ai/)
  • Join the AI Facebook Community: [AI Facebook Group](https://www.facebook.com/groups/739308654562189)
  • Podcast Studio AZ: [Podcast Studio AZ](https://podcaststudio.com/mesa-studio/)
  • Podcast Studio Network: [Podcast Studio Network](https://podcaststudio.com/)

Privacy Policy

See Privacy Policy

[Privacy Policy](https://art19.com/privacy) and California Privacy Notice: [California Privacy Notice](https://art19.com/privacy#do-not-sell-my-info).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The wait is over. Dive into Audible's most anticipated collection, The Best of 2025. featuring top audiobooks, podcasts, and originals across all genres. Our editors have carefully curated this year's must-listens from brilliant hidden gems to the buzziest new releases. Every title in this collection has earned its spot. This is your go-to for the absolute best in 2025 audio entertainment. Whether you love thrillers, romance, or nonfiction, your next favorite listen awaits. Discover why there's more to imagine when you listen at audible.com slash best of the year. The whole field of text generated AI is, of course, familiar to like everyone.

0:45We all use ChatGPT. I think when you combine that with images, we're starting to see a lot of really exciting things. Once you now have like Dolly built into ChatGPT, for example, there's a whole bunch of new applications that I've been testing out and trying. And I think the next step in this whole AI wave right now is going to be really video. And I think that AI generated video and using AI in the video field is going to be absolutely huge. You know, there's already a lot out there, but I think this is going to grow and become very mainstream. So an example of this is 12 Labs, which is a startup based in San Francisco, and they are pioneering the field of, I think, video language alignment.

1:23So 12 Labs essentially is trying to provide a robust framework for multi-modal video understanding. and their initial focus is going to be on semantic search. So Jay Lee, who's a co-founder and CEO, likens it to, you know, control F for videos. I think that's a hilarious description of the company, which I'm sure is probably command F on Macs. But in any case, he recently said, quote, our vision is to equip developers with tools to create programs that perceive, hear, and comprehend the world just like humans do. So by linking natural language to the contents of a video, 12 Labs models identify actions, objects, and even background noises.

2:07And this offers developers the opportunity to design applications that can sift through videos, categorize scenes, extract key topics, break videos down into chapters automatically, and a lot more. I think this is absolutely fascinating and really critical. um i talk a lot about like really generative video ai is going to be huge we're seeing companies like runway kind of tackle this although like i think runway is a cool example because they recently did like a big um like a video production competition thing or like you know uh award or like uh yeah pretty much academy awards thing but with generative ai i saw a lot of the videos that were generated for that and like it was cool and impressive but like obviously they weren't super high quality, like this is generative AI, it's not perfect.

2:53So if you if you messed around with like, generative video before on runway, pretty much you pretty much know what you're working with, I think this is going to get a lot better. And in order for it to get a lot better, we're going to need companies like 12 labs to really build out the infrastructure around this. And so something that I think that's fascinating that they're doing here is they're essentially building AI, they're using natural language to understand what's happening in a video really understand the deep context. And that I think is the first step to being able to essentially use generative AI to actually create video.

3:25Video is so fascinating because it's a combination of so many different other elements and areas of AI. There's obviously the natural language processing from ChatGPT. There is the image generation and stable diffusion aspect that, you know, mid-journey is kind of pioneering or, you know, Dolly 3, whatever. And then you have to bring it to life and if you think about it like a video obviously is just a whole bunch of frames of images so you understand the concept right it's like obviously you tell it to do something but in order to like create a frame of video it's got to create like you know a couple hundred generative ai images and so it just is a fascinating concept we're really bringing together and then of course there's the audio aspect of you know grabbing something like what 11 labs is doing to create voices and audio and then great creating other ones that are going to create like background sounds like if you if you go to uh you know runway eventually and say hey generate me a scene that's like a beach with a seagull and a family talking and the dad in the family saying hey we all need to get into the car a storm is coming like there is so much that's going to go into that scene it's like you have to generate all of the images that are going to create the family doing the movements um you have to make them all the same obviously you have to you know use chat gbt and natural language processing to come up with the script and exactly how they he says that and how people kind of might respond or react then you have to have like beyond just the audio aspect you have to have the sound effects in the background you have to have the birds chirping and the waves crashing and the waves crashing has to line up with the actual waves crashing in the video like there is so much that's going to go into this and i don't doubt we're going to have this nailed down in the next one to two years up to a very high level but it is a big challenge.

5:05I think what 12 labs is doing by essentially allowing developers to understand very deeply what's going on in a video, then they essentially reverse engineer that to create videos. So I think this is kind of the technology that's going to bring us there. So I'm really excited to see this and kind of cover this. But I think talking a little bit about the utility of 12 labs technology, Lee also, their CEO of 12 labs mentioned its potential in ad placement, content moderation and also media analytics a prime application i think would also be in discerning the context um in which you know knives appear in videos you know whether they are showcased in a violent scenario or for educational purposes or like you know i think you could say like you know generate a scene in a kitchen where there's like a knife you know that's like falls down and cuts something and it's like the the gender vi has to know is a person holding that knife cutting something?

6:00Is a person dangerous? Is this knife haphazardly fall off the counter? There's so much context that needs to be given. And so I think, yeah, really having tools like this that look at video, look at what is common, and then they're able to essentially use that to generate video in the future, I think is going to be important. So the technology right now can autonomously generate video summaries like blog posts, headlines, or tags. I think drawing a distinction between their offering and large language models like chat gbt lee who's the ceo of 12 labs uh kind of emphasized that their solution is engineered to understand videos a little bit more holistically um integrating visuals audio speech elements right so really talking about all the different elements that are in a video he said quote our focus has always been on pushing the boundaries of video understanding there's a bunch of big tech giants like google that are also kind of getting into this space with models like MUM, which essentially powers video suggestions on, you know, Google search and YouTube.

7:03So Lee believes though, that 12 labs stands apart due to essentially its model quality and also customization features, enabling clients to modify the models for specific video analysis needs. In recent developments, 12 labs has unveiled Pegasus 1, which is a state-of-the-art multimodal model equipped for a wide array of video analysts tasks. So Lee kind of underscores the importance of these advancements saying, quote, conventional video AI models lack the depth required for many business applications. Our goal is to provide enterprises with tools that achieve human-level video analysis, eliminating the need for manual scrutiny.

7:45I think since its private beta in May 12 labs has amassed a user base of 17 000 developers so they have a lot of people using them the clientele is across a bunch of different sectors from sports and media to e-learning and security there's a bunch of different you know prominent names like the nfl um of course in their investment pool i think they just raised 10 million dollars um and this came from nvidia intel and samsung next um and that takes their total capital raise to around 27 million so they're a fairly serious player in this space. Lee sees the influx as a catalyst for innovation and growth.

8:21And he said, quote, our objective is to further the field of video understanding and present the best models to our clients, allowing them to achieve remarkable feats. I think this is a really interesting space. I don't know. The fact that the name of the company is 12 labs and my favorite audio producer is 11 labs. Now I got like 11 labs for audio and 12 labs for video. I I don't know. It's funny. I don't know if that's like super creative or what the what the thinking behind those names are. I guess I got to get the CEOs on and ask them. But in any case, I think this is a big space. Video is the next frontier that is going to be conquered by AI.

8:56And I think 12 Labs is going to play a big role in that or companies like 12 Labs. So I'm excited to follow along and see how that actually rolls out. Really, really excited for generative AI and video. I know it's around the corner, so stick around and I'll keep you updated on all the advancements and let you know when this is a serious technology that you can try out.

From the publisher

In this episode, we explore the recent funding announcement by Twelve Labs, discussing how the $27 million investment will fuel the next wave of AI in video and the potential impact on various industries.

See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

More from AI Today

All 897 episodes
Twelve Labs Secures $27M Funding for Advancing AI in VideoAI Today · 9 min
Listen in VO