In short
TWIML AI Podcast Episode #660: Data, Systems and ML for Visual Understanding with Cody Coleman
Episode Overview In this episode of *The TWIML AI Podcast*, host Sam Charrington interviews Cody Coleman, co-founder and CEO of Coactive AI. The discussion revolves around Coactive's innovative multimodal asset platform and visual search tools, as well as key machine learning (ML) methodologies such as active learning and core set selection that enhance the efficiency of the machine learning lifecycle.
Key Guest
Cody Coleman
- Position: Co-founder and CEO of Coactive AI
- Background: PhD from Stanford University focusing on democratizing machine learning and artificial intelligence.
- Core interests: Data-centric AI, multimodal embeddings, and infrastructure optimization for scaling ML systems.
Episode Highlights
- Background and Journey to Coactive AI
- Cody's academic focus was on lowering barriers in AI, particularly through resource and data-efficient deep learning.
- Developed the DawnBench benchmark to track ML system performance, which later evolved into MLPerf and MLCommons aimed at democratizing AI.
- Understanding Active Learning and Core Set Selection
- Active learning is crucial in reducing data labeling costs and improving the efficiency of ML workflows, despite being underutilized in academic research.
- Core set selection helps optimize which data points to label, mitigating the need for excessive data.
- Coactive's Multimodal Asset Platform
- Coactive is described as a multimodal asset platform (MAP) designed to enable easy content searching and analysis, flipping traditional data processing methods (tag load search) to loading and searching raw content without tagging.
- This approach improves flexibility, speed, and cost-efficiency, similar to the transition from ETL to ELT in data processing.
- Comparative Analysis with Other Platforms
- Unlike other tools focused solely on computer vision or ML engineering, Coactive targets media and asset managers, enabling them to leverage content without deep tech knowledge.
- The platform addresses the need for structured workflows around unstructured visual data, making it accessible to a broader audience.
- Technical Underpinnings of Coactive
- Coactive’s infrastructure is built to handle vast amounts of unstructured data—images, videos, etc.—by optimizing data locality and ensuring efficient embedding processes.
- The platform’s architecture allows for rapid re-embedding with new models, keeping systems future-proof amidst fast-evolving AI technologies.
- Productivity and User Experience
- Coactive enhances productivity by automating tagging through techniques like dynamic tagging, which leverages active learning to reduce manual labeling efforts significantly.
- Real-world example: Fandom, a large fan-generated media platform, dramatically reduced its manual content review process using Coactive’s tagging system, achieving 85% efficiency.
- Lessons and Core Values for Startups
Cody emphasizes several guiding principles for building effective AI products:
- Simplicity: Transition from complex, waterfall ML processes to agile methodologies.
- Scale: Solutions must be able to handle vast oceans of data, especially as unstructured data grows.
- Security: Strong focus on data governance and security is essential.
- Model Agnosticism: Flexibility in model deployment to adapt to rapidly changing AI technologies.
- Cloud Agnosticism: Ability to integrate with various data sources and remain adaptable.
Conclusion Cody shares insights on the transformative potential of AI and emphasizes the continuous evolution of generative technologies. He highlights how Coactive is positioned to democratize access to visual data analytics, supporting enterprises in navigating the complexities of modern AI applications.
Episode Resources For more details on this episode, visit the complete show notes at [twimlai.com/go/660](https://twimlai.com/go/660).
Key Takeaways
- The significant role of data-centric methods in improving ML efficiency.
- The necessity of multimodal platforms in modern data processing and analytics.
- Practical advice for entrepreneurs on building scalable, secure, and flexible AI solutions.
Final Notes This episode highlights the intersection of advanced AI methodologies and practical applications in enterprise settings, underlining the importance of democratizing access to sophisticated tools for a wider audience.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:09All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Cody Coleman. Cody is co-founder and CEO of Coactive AI. We are, of course, coming to you live from the Future Frequency podcast studio here at AWS reInvent Conference, which I've been covering via X and LinkedIn. Be sure to find and follow me there for the latest reInvent and AI updates and insights. Cody, welcome to the podcast. Thanks for having me. It's awesome to be here, Sam. I'm super excited to have this conversation with you, especially since before we started rolling, you talked about how you listened to the podcast when you were doing your PhD.
0:50And that was Stanford? Yeah, Stanford University. So I think I started right around the time that the podcast did back in 2016. That is awesome. It's just been awesome. You know, like I spent my entire career at the intersection of data, systems, and machine learning. So your podcast and the topics that you cover really, really resonate with me. Very cool. So tell us a little bit about your background and what you do your PhD in and what brought you to founding Coactive. So I spent all of my professional and academic career at the intersection of data, systems, and machine learning. And back in 2016, when I started my PhD, I joined the Dawn Project at Stanford, which was about democratizing machine learning and artificial intelligence.
1:32Because the thing that I loved about computer science was the fact that all you needed, like when I was growing up was a computer and an internet connection to create something that would impact the lives of thousands, millions, or billions of people around the world. But with AI, you know, it started to change. There was all these barriers. You needed like a tremendous amount of compute in order to be able to do anything, a tremendous amount of data to do anything, and a tremendous amount of expertise. So I focused on bringing down those barriers. Like my dissertation was resource and data efficient deep learning.
2:01Okay. And how did you approach that in your dissertation? So the first part of my dissertation, thinking about the computational resources aspect of it, I created the first end-to-end benchmark focused on ML system performance, DawnBench, which really got the industry to focus on training time and training costs and inference costs and inference latency. And then that grew into MLPerf. And MLCommons. And MLCommons, yeah. It was kind of a wild journey. Second year as a PhD student, I'm like, who's going to pay attention to me with releasing DawnBench? And then just like years following after that to see that grow into MLPerf and then now MLCommons, this whole nonprofit, really at that mission of democratizing AI to make sure that it benefits everyone.
2:43Yeah, awesome. And so from the PhD to Coactive, what was that path? So the first part of my PhD focused on the computational resources piece of it. And then the second part, you know, is thinking about bringing down the data barrier. So I was thinking about, you know, there had to be a smarter way than just tossing like every data that you can think of at like a machine learning model. So I did research into active learning and core set selection to be smarter about what data points we train on and that we label. And through that experience, I was able to work at leading tech companies like Pinterest and Meta and seeing how they were actually able to leverage AI to work with all their content and improve things like search, ads, recommendation, protect the safety of online communities, protect copyright material.
3:29And that there was this kind of emerging playbook that was forming around AI, but there was no enterprise grade solution for doing that. And you needed to have like a bunch of PhDs like me to be able to do anything. I saw that as kind of an opportunity to build something here and that there was a real need. And that was the genesis for Coactive. Active learning always struck me as an underutilized technology. I'm wondering if you share that perspective or if you saw it in use at all those places that you mentioned. You're definitely preaching to the choir here. I'm biased in this regard. Active learning is such a funny technology because from an academic perspective, It's kind of a neglected area of research when we think about kind of everything that's happened and kind of modern artificial intelligence machine learning.
4:13But when we go into practice, because of the fact that it's so expensive to label data and so costly and slow, all the big tech companies were doing active learning in order to bring down costs and to make things faster and more accurate. And that's only gotten more and more important as we have these kind of models that can get us general understanding 80 % of the way there. And now it's really kind of the key challenge around machine learning is figuring out what are those right data points, those few examples that you need in order to be able to fine tune to your specific use case, where active learning is more important than ever in practice.
4:44Yeah. But from an academic perspective, it's understudied. So it sounds like you're saying it is used pretty broadly in practice, only at larger companies, or has it trickled down to smaller companies? I know last year, a couple of years ago, maybe, we really covered on the podcast the idea that Andrew Ng popularized data-centric AI, which seemed to kind of tug at some of the same strings, but not necessarily centered on active learning as an approach. Yeah, it's actually so funny that you bring that up. We actually co-organized the first workshop on data-centric AI at Neurots with Andrew Ng and folks from Meta and Google back then.
5:24And with that whole kind of movement from model-centric AI to data-centric AI, active learning and core set selection are kind of fundamental technology, a part of this broader movement to focus in on data. As we've seen models kind of, of course, there's always progress, but it's standardizing in terms of like transformers taking over the world and it's focused on data. So the data-centric AI movement and focusing on things like data quality, data selection, data cleaning, like week supervision, all these things, very much in line with the research I was doing at Stanford. And I was excited to be a part of creating that first workshop at NeurIPS on data-centric AI.
6:01And then also in working with ML Commons as they created the Data Perf Benchmark Suite for data-centric AI. And we even helped shape the vision data selection benchmark in the Dataperf benchmark suite. Okay, awesome. Now, as much as we're talking about kind of low-level infrastructure and platform stuff, and that was the environment that you kind of grew up in career-wise, Coactive isn't doing, it's not a tool play, it's not a platform play, it's more of an application. Is that right? What's the right way to think about it? So Coactive is, as I would describe it, a multimodal asset platform, or MAP for short, that makes it easy to search and analyze content.
6:42And kind of what I realized very quickly after I left academia and started Coactive and started working with enterprises is just realizing that there's so much infrastructure systems and processes that need to be like stitched together from like an infrastructure platform perspective to do any of this. And that for many companies, for many enterprises, it's a huge barrier that they can't overcome. Taking a step back, when you think about how enterprises work with image and video data, before Coactive and before this recent wave of multimodal learning, you had to do this tag load search process.
7:20You know, you had to tag the raw images and videos, either through human or machine annotations. And then you would load those annotations into your systems. And those annotations are nothing more than like a JSON file with a bunch of words. And you would search based off of those annotations. And, you know, that's a slow, expensive, and inflexible process. And fundamentally, what we're doing at Coactive is we're leveraging AI to flip that on its head with a load search tag approach, where we can actually load and index the raw images and videos and make them searchable without any metadata or tags whatsoever.
7:54And then... Is this an explicit nod to kind of the shift from ETL to ELT? So like I'm thinking about tag load search TLS to load search tag LST world, you know, I think of it very akin to the transition from, you know, ETL to ELT, you know, where effectively you get greater flexibility, at greater speed, and it's just a more agile and cost-effective way in order to be able to work with content. Just in the same way that in the data warehousing world, doing ELT can give you greater flexibility and allow you to do transforms after the data has already been loaded. What you're describing in terms of a kind of a multimedia, multimodal asset platform reminds me of, I'm forgetting the name of the individual and the company, but also at Stanford.
8:40I know you know who I'm talking about because he ran a systems conference out of Stanford for a while. And he had a video platform company, Matroid, I think. Oh, yeah, yeah, yeah. Reza? Reza, yes. Are you going after similar ideas from a product perspective? Yeah, so super great question. You know, fundamentally, when I think about it, we're trying to make it easy to search and analyze content. Whereas I would say like there's a lot of people that are, you know, their mission is more make computer vision easy. And there's a subtle but important difference there, you know, in terms of like the audience and in terms of like the problems that you're solving.
9:20There's a massive ML community out there and making like the process of doing computer vision easier, providing that infrastructure is great and super valuable and it targets that ML audience. But fundamentally, with Coactive, we wanted to target kind of the more traditional data audience. Thinking about this like past decade in the big data movement, we've established all these kind of great workflows around ad search and recommendation centered really around structured and semi-structured data, you know, tables and documents. And this massive ocean of unstructured visual content has sat kind of outside of that world.
9:56And fundamentally, we want to bring structure to the unstructured data so that it can fit into this big data movement and the workflows that people already have today. So as opposed to targeting the tool at a computer vision engineer who's trying to analyze lots of video and wants something to accelerate that work, you're targeting, you're offering at media managers or asset managers at an ad agency or at a large brand who don't know or care anything about computer vision, multimodal models, that kind of thing. Yes, exactly. Our platform caters to both technical and non-technical folks within enterprises.
10:33Because I know for like you and I, like AI has been like has seemed ubiquitous for like, you know, years now. But for a lot of enterprises out there, you know, outside of like Silicon Valley. It just started exactly one year ago on November 30th, right? Exactly, exactly. You know, we're still like doing like data transformation, digital transformation and things like that. And just getting that past decade of big data and now having to deal with AI and try to create new systems while they're still thinking about just big data processes and things. We wanted to target those folks, again, kind of coming back to that mission of democratizing AI.
11:10So it's not just the tech-first companies that have a tremendous amount of PhDs and ML engineers that can benefit from this next wave of intelligent applications. that it's all the other organizations as well. And that's even individuals. You know, I was going to say like you and I, but we're probably not the average person out there. But really enabling everyone from like, you know, media and entertainment companies that have to get content out of the door faster to monetize it on streaming or social media. We have editorial teams, you have marketing teams, you have production teams, where their success depends on their ability to be able to search, filter, and analyze content.
11:48In consumer retail, when you think about the rise of e-commerce, we're making purchasing decisions based off of images and videos and the systems haven't really adapted. You know, kind of one anecdote that I love here is that there's a large fashion company and they did this massive marketing campaign around what, you know, you and I would probably call ripped jeans. And the, you know, the marketing campaign was successful. You know, it drove a lot of traffic to their website, a lot of searches. But the problem was that when users search for ripped jeans, nothing came up. And the problem was that when they looked into the data, when they looked into the SKUs and the products that had ripped jeans, they were labeled as distressed jeans or tagged as distressed jeans.
12:33And that disconnect, you know, as soon as they fix that and they change from distressed jeans to ripped jeans, those products sold out like that. Yeah, interesting, interesting. Yeah, one thing that I've been saying for many years now is that search sucks. It's just hard. And I've been kind of revisiting this in the context of RAG, like everyone is talking about, hey, let's just throw a vector database in front of an LLM and now we're going to get all these wonderful responses. And there's still a lot of hard work that has to go into optimizing that retrieval. And search folks have been trying to get that working for many, many years.
13:12It's not easy. That's what the example you described kind of reminds me of just the difficulty of getting search type experiences correct. And now these dialogue experiences are actually search experiences under the covers, right? Yeah, exactly. It's funny, search seemed kind of stable for the past two decades, in a sense. And it's such an exciting time now when we think about these foundation models and everything that's happening in information retrieval as a result of that. We're now, everything that we know about search is kind of like, it's really a paradigm shift in how we think about searching all forms of content.
13:49Meaning from kind of a keyword-oriented paradigm to more of an embedding or vector-oriented, vector search-oriented paradigm. Exactly. Where it's like rather than this kind of like having to have discrete like labels being able to actually effectively calculate the DNA of content. A vector embedding is just a list of a few hundred or thousand floating point numbers, but it captures all the semantic information that's captured in an image, a video or an audio file. In the same way that, you know, you can pull like a hair off of your head and you can get like a strand of DNA that captures everything that describes your genetic makeup and who you are.
14:26We're having this DNA moment. I remember when DNA was first sequenced and it was a huge thing. And everything that fell out of it. And it's such a fundamental paradigm shift. Just in the way that editing DNA to create something new like CRISPR is a massive thing. That's kind of what's happening in generative AI when you think about being able to actually edit and use these embeddings to generate something new like in a RAG type of setup. And in the same way, you can also use embeddings to diagnose problems, just in the same way that, you know, you look at something like 23andMe and you can use DNA to diagnose diseases.
15:01And you can do it for information retrieval to search a massive database of content, which is incredible. Let's go a little bit deeper into Coactive and kind of the technical underpinnings that enable you to do what you do. I'm imagining that embeddings and as a part of that, we heard in the keynote this morning and Swami's keynote here at reInvent. He actually spent a bit of time talking about the complexity of embeddings and in particular, multimodal embeddings. Is that part of what you're tackling? It's a part of it. And when you think about it, it kind of goes back to this old paper about the hidden technical debt of like ML, you know, where it's like the embedding in the model is like, you know, kind of one.
15:40It's a very important piece, but it's only one piece of the overall system. And when I think about what we built at Coactive, it's effectively like we've built a car. We've built out kind of all the systems, all the infrastructure. And then we can swap out. We take a model agnostic approach where we can swap out different encoders, different like embedding models as easily as it would be to swap out tires on a car. Because ultimately, you know, the thing that I think we've seen in this past year is that AI is moving so rapidly. And it's unclear what's going to happen in the next 12, 18 months as far as what model is going to be best or anything like that.
16:16But companies, enterprises have to build today. By creating this platform, this multimodal asset platform, people can build today but future-proof themselves as we just continue to see rapid progress in the modeling piece of it. One other thing that I wanted to touch on, it's actually quite interesting because you mentioned this around search where, you know, the way that I think about it is that there's like AI people that are trying to reinvent and like rediscover everything that we've like learned in data systems and databases, you know? Information retrieval. Exactly, exactly. Information retrieval, you know, scalability.
16:50Distributed computing. Distributed computing, data providence, data rights, data privacy, data governance, all these things. Not to mention the humanities and ethics. Exactly. So you have this one side from the AI side kind of rediscovering everything in databases. And then there's also this same thing from the Swami's keynote. You have all the databases kind of piece of it actually trying to integrate the AI approach as well. And fundamentally with Coactive, what we're doing is we're taking an AI-first approach to data systems where the two things are really married together rather than from the very beginning, rather than kind of like AI over here and databases over there.
17:28Yeah, what I would love to get at kind of in prompting you to go a little bit deeper is like, you've built this system that kind of targets a non-necessarily technical end user or both that uses, you know, it's kind of cutting edge in the sense of like, we're trying to figure out multimodal AI. You're doing multimodal AI. What are the lessons that you've learned or the techniques that you've stumbled across or the cool things that you figured out that other folks that are building kind of in the space, you know, multimodal or similar, using similar types of technologies? What are the things that you've learned that others can benefit from?
18:07Understanding systems is super important. I mean, every kind of breakthrough in AI has kind of been around systems. And, you know, maybe taking a step back to kind of put it into perspective, when we think about the past decade and the big data movement, it's focused primarily on structured data and tables. And if you think about 10 million rows of tabular data, you know, that's about 40 megabytes. If you think about 10 million documents, now you're talking like 10 million pages from Wikipedia. Now you're talking about 40 gigabytes. That's three orders of magnitude different. That's like going from the surface area of Lake Tahoe to the surface area of the Caspian Sea.
18:44When you think about 10 million images, let alone like video, 10 million images is about 20 terabytes if you look at the open images data set. That's another three orders of magnitude difference. And that's like the surface area of the Pacific Ocean. So when we think about kind of the tools that we've built today for, you know, the big data movement, they're great for working at, you know, the scale of like a data lake. You know, it's kind of like having like a rowboat or a canoe. Like, it'll be fine to get across the lake. But if you told me to cross the Pacific Ocean with rowboat, I would think that you're crazy.
19:16And that fundamentally, you need a bigger boat. And that's kind of core to how we think about what we built at Coactive. And then with that high-level context that kind of out of the way, you know, I can go through kind of piece-by-piece things. So first off, when you think about where does data sit? You know, where do images and videos sit? And it's in S3 as a bunch of individual files and a network file system in S3. So first, a system's perspective, rather than trying to read a bunch of small files from a network system, coalescing that together into some form of a binary format, whether it be Parquet, LMDB, things like that, will give you a dramatic improvement in performance in your downstream systems to make things way faster.
20:02And then from there, you can embed very quickly. And then you can also re-embed as new models come out and there's new improvements and try things out way more quickly because you have all the files coalesced together in a single thing. So thinking about data locality, which is kind of a core piece of systems, is really, really huge at that very beginning piece. And then something that Swami mentioned earlier today as well is thinking about all of your data kind of coming together. When you think about images, you have an object store, you have a relational database that has information and metadata about it, then you have this vector database as well.
20:39When we think about actually managing the embeddings to actually make them searchable and to be able to do other things on top of it. So thinking about how do you keep things consistent across all these different forms of effectively the same asset, a tough challenge, but something that I think is really important in order to be able to provide the interface that people have known around databases and working with data. But then once you do that, once you do that heavy lifting, which is kind of like part of the core systems piece at Coactive, it then gives you this unifying layer at the embedding kind of stage where you can do all sorts of things on top of it.
21:15So we can do search, like intelligent search is kind of our first capability. But then we have dynamic tagging where we can actually leverage active learning and leverage data-centric AI techniques to make it very quick And, you know, almost like going from a waterfall approach for ML development to an agile approach to ML development, where interactively you can define and fine tune your models and define classifiers for domain specific concepts and then generate consistent metadata over your entire catalog of content. What's an example of that? Can you kind of make that more concrete? Yeah. So I'll talk about fandom here.
21:50So Fandom is the world's largest fan-generated entertainment and gaming platform. They've been a champion and the end-to-end resource for 350 million fans worldwide. And their users every year upload tens of millions of images to their platform. And by and large, a lot of that data is really good and really valuable. there's a small fraction of it that violates their community guidelines and corrupts the safety of the communities, the fandoms that they've created. What they had to do previously in order to be able to kind of like their community guidelines and codifying their community guidelines or to enforce their community guidelines, they used to have human beings review every single image that was uploaded to the fandom platform.
22:39You know, kind of going back to this tag load search mentality. You know, you had to tag the raw images and videos in order to be able to do anything with it. So they were very much in that tag load search world. And this is super problematic for a lot of like reasons, you know, one, it's expensive actually having human beings reveal tens of millions of images every single year. It's also slow. You had to wait like 24 hours in terms of an SLA for a human being to go through that review process. And even thinking about ethics, there's just some content out there in the world that like someone's got to look at it or that really should never see the light of day.
23:12You know, it's like, unfortunately, these people have to like, you know, look at the the darkest parts of the internet. And with Coactive and working with Fandom and doing dynamic tagging, we were able to actually codify, enable their community team, who's a non-technical team, to take their community guidelines, take the process that they have been doing, codify that into these dynamic tags, and then generate metadata over every asset that they had in the past, as well as all the new assets that are coming in. And what we were able to do is reduce their manual labeling efforts by 85%. And originally, we like to over-deliver.
23:51We were like, hey, it's going to take you six months to be able to recoup the value of this. They saw the kind of performance and cost savings results that we have projected or predicted for six months in two months. So way faster in dramatically reducing the amount of data that they actually have to go through and manually annotate by generating that consistent metadata over all of their content. And that's kind of just like just a day one problem. The real powerful thing, they had been looking, actively searching for solutions like AI powered solutions to do this, but no existing solutions would solve their problem because of the fact that they were kind of one size fits all.
24:25They weren't able to be customized to that last mile for different fandoms. Like you might have a different set of community guidelines for the fandom around Game of Thrones than, you know, Puppy Patrol. And being able to actually kind of codify that and understand specifically for their use case was huge for them and being able to actually, you know, take the, save people from having to look at like the worst content that's out there. Now, the process you described calls to mind ideas or techniques like weak supervision, programmatic labeling, are some of those, are those actual techniques that you're using as part of delivering the solution?
25:00So weak supervision in a sense, since weak supervision or semi-supervised learning is like a very broad, you know, category of techniques. So in that regard, we do stuff in like weak supervision and semi-supervision. The programmatic labeling kind of piece of it is a little bit more difficult because when you think about, you know, labeling functions and things like that. Plain those to images and media is like the classic define a cat in this picture problem, right? Exactly. You know, like text is like almost like a discrete problem. You know, you have this whole vocabulary, you have these like tokens of words and things like that.
Read the full transcript
25:35So you can, like a human being can be like, you know, if blank is married to blank, you know, you can write like a labeling function and a rule to be able to kind of predict that based off of text. But, you know, I'm not an artist. And if you asked me to describe a bunch of labeling functions for like a cat or like, I don't know, a bag of chips or something like that, like it would be impossible for me to do. So fundamentally, that's where being able to do stuff like active learning as kind of a starting point where you can point to actual visual examples really helps to be able to actually kind of quickly get that signal and then focus in on that signal to be able to kind of very accurately, very quickly provide predictions.
26:13Because, you know, text is, again, kind of like a discrete problem. And then visual content is more like a continuous problem, you know, going back to like my signal processing days, where you're looking at like the level of like raw pixel values. Yeah, interesting. So then that kind of reinforces the need to have a very solid foundation of embeddings and things like that. It sounds like you're relying heavily on that for providing this functionality, like identifying neighbors of some image that was identified and using that as a way to do the labeling. Yeah, to dive in there about like, how do how do embeddings fit into active learning, right?
26:54Because it's interesting, like the active learning literature, you know, was kind of started before this big data movement. So it focused on, you know, data sets of like a smaller scale, like a few tens of thousands, you know, hundreds of thousands was like large scale active learning in the literature. And you could do things that were quadratic, and you could search over the entirety, like all of the examples. But when you think about kind of this big data movement where we have millions, billions of data points, it's impossible for you to go to all the raw assets for each individual loop of active learning in order to do it.
27:29It's just not performing and scalable enough for this big data eras. So that was actually part of the research that I did during my PhD. I was working at Meta, published this paper, similarity search for efficient active learning and search of rare concepts. And the core idea there was the fact that we could, you know, these large language models, these foundation models, you know, if you think of the GPT-2 paper, it was called Large Language Models are Few-Shot Learners. And the key kind of innovation there is that we have these kind of generic representations that actually can cluster unseen concepts pretty well, like, together in this latent space.
28:02So then rather than doing this global search over all of, like, your unlabeled data to find the most informative data points, you can instead, in the embedding space, start locally and expand based off of that. And what that does is it goes from like, you know, I remember doing experiments where it took like more than a day to do a single round of active learning over like 10 billion images to something that could be done interactively in like a Jupyter notebook on like my dev machine. Meaning because you're no longer working with these large assets, you're working with much more dense representations of those assets.
28:40Exactly. Very cool. Yeah. And it's kind of funny. I love the dense representations. It takes me back. It goes back to like word to VEC, you know, it's like when embeddings like first happen, you know, like the history of embeddings. I know that like vectors and everything like that are really popular now. But I mean, it's been interesting to see this kind of whole evolution over like the past decade around embeddings and dense representations. And so are you specifically doing multimodal embeddings as part of your approach? Yeah, we're specifically using multimodal embeddings. So, you know, and multimodal is a little bit of a kind of a suitcase word in kind of my mind, because I mean, the way that I think about like multimodal is it's about kind of the modalities of the inputs to the modalities of the outputs.
29:23Like stability AI is kind of multimodal in the sense that you start with, or stable diffusion is like multimodal in the sense that you start with like one modality, text, and then you generate like the output is a different modality. It's visual. So that's like one type of multimodal in a sense. But then you also have kind of like what you see in kind of maybe more of the large language model kind of multimodal space where they'll have, you know, image and like text as inputs. But then the output is just text. So that's multimodal as well. And fundamentally, when we think about the multimodal embeddings that we use at Coactive, your input might be like a text prompt or text and image.
30:01but then the output is going to be images, finding the relevant images that match that query or being able to say, you know, is this image kind of associated with this general class or things like that based off of the embedding. So we use multimodal embeddings in the sense of going from image text and images or just text by itself to images. And so we've talked about kind of the system platform. We've talked a little bit about kind of the embedding aspect and how you're using that to enable kind of this automated labeling capability. What other interesting bits or learnings are there that folks can take away from, you know, what you've built in the approach?
30:43Yeah, thinking about kind of, so Coactive is two and a half years old. And just thinking about kind of as we engage, you know, as we have engaged with enterprises and more and more customers and things like that, why people choose Coactive kind of comes down to like five core things. And I think these are lessons that any practitioner or any founder or researcher can learn from. You know, first is simplicity. Going from this waterfall approach of ML development to a much more agile approach of ML development. Because when you look at like prior to like embedding-based things and things like RAG, where you can actually iterate very quickly, it was very much like a waterfall thing.
31:22You had to have like a data engineer, get your data into a good place. then an ML researcher would come in and figure out what model you would train. Then you would have to label data so there'd be an annotation team. Then it would go back to the ML researcher to train the model, then MLOps. And then finally, after all that, it gets to business intelligence. What we're saying with embedding-based things is that there's kind of this great decoupling that's happening from the slow, expensive part of deep learning and AI. Meaning someone else is doing that, the models are pre-trained, you don't have to think about it.
31:51Exactly. And then instead, you can decouple that from all the different downstream tasks that you, as a business, as an enterprise, care about. And I think of it from a system's perspective as it's almost like an embedding is a cache for computation. You're basically, rather than having to pass through all those layers, all the hidden layers in a deep neural network, over and over again every time you process an individual asset or piece of content, you just have the embedding vector. And that cache is that computation. So unless you're trying to change the relationship between the entities themselves, you can just leave that alone and have that be a proxy for those assets.
32:28Exactly. And that enables this kind of much faster kind of loop and process. And then the second kind of piece of it is just scale, kind of going back to like 10 million. Like we have to move from data lakes to data oceans and to really be able to unlock all of the data that we have in the world. Goes back to this old idea, like Bill Gates said, that content is king. And he predicted in 1996 that the real money on the internet was going to be made in content, just as it was during the broadcasting era. And that, you know, the real long-term winners were going to be the people that were able to effectively leverage their content to deliver information and entertainment.
33:03And when you fast forward to today, those predictions have come true and the king is here. 80 % of internet data today, or 80 % of internet traffic today is video data. And it's predicted that by 2025, 80 % of data worldwide is going to be unstructured data, so it's just audio and video files. And that prediction was made before this recent wave of generative AI, which has dramatically lowered the barriers. So, you know, it's like if it was a tidal wave before, it's a tsunami now. And it's just a massive amount of content at like a different scale than what we've seen before. You know, 80 % of data out there.
33:32When we think about the systems that we have, like it's really just like the tip of the iceberg in order to process that. So being able to scale with this like massive ocean of data and move from a data lake to a data ocean is huge. And then security. like enterprises i think as we've all seen over the past year like data rights and security data governance is a huge thing and ensuring that your data stays your data is almost like table stakes for a lot of enterprises and then there's the model agnostic piece that i mentioned as well which is over this past year every single company every single enterprise that's out there has a mandate to like ai is an existential threat to their business like if they don't use ai if they don't build things today, they might be left behind.
34:10But at the same time, AI is in its infancy. You know, we're evolving very quickly. Exactly. Very, very quickly. You know, it's like one week company A is like, you know, leading and then they're like, they have the best model. Then company B comes up and then like, it's like, oh, you know, actually open source is like out there and like amazing. And then another, and it's just like, and company A becomes a soap opera for 10 days. Yeah, exactly. Exactly. Exactly. You know, at least a lot of interesting discussions and things like that, you know, it definitely keeps you entertained over the holidays.
34:42But like, fundamentally, since we're like in the infancy of like this AI era, you know, we're in like, when you think about the big data movement, we're like 10 years into the big data movement. This is going to be another 10 year journey or more. And we're at year one right now. And it's going to evolve so much. I mean, you see that today here at like AWS reInvent, just the sheer number of foundation models are coming out, the constant improvements there and being agnostic. so you can build today and future-proof yourself for tomorrow is huge. The other thing is being cloud-agnostic, being able to actually read data from wherever it is, you know, because data is everywhere.
35:15There's a lot of dark data out there, especially when we think about unstructured data that is just sitting all over the place and being able to bring that into one place and unify it across all these different data sources is really, really critical to actually being able to have to make sense of all of it. And those five reasons, simplicity, speed, and scalability, security, model agnostic, and cloud agnostic are kind of the core kind of lessons and the core reasons why companies choose Coactive today and something that any practitioner or startup or enterprise. And no matter what you're building.
35:49Yeah. To keep in mind, it's like, it's huge. Awesome. Well, Cody, thanks so much for taking some time out of your busy reInvent schedule and saving some of your voice for us. Yeah. It's been great to chat. Do you have something to throw in? Oh, yeah. One thing that I did want to add is, you know, we're at reInvent. I just also want to say, like, it's been interesting. We're two and a half years old as a company. And we're like, you know, we're like 20, 25 people now as like a company. And AWS has been like an amazing partner. I remember when it was just like my co-founder, Will, and I, we started the company and we're like, oh, my gosh, we're like two guys and like an idea, like how we're going to do it all.
36:24And like kind of jumping out into the wilderness and like the fear of being alone. But with AWS, we've never been alone. Like I remember it was like 2 a.m. We were doing like a large ingestion job and we were like reaching like our quota limits and things like that and getting alerts. And it was like 2 a.m. and we, you know, texted our account team and being like, hey, can we get a quota increase at 2 a.m. so that we can actually meet this like deadline that we had for a customer? And, you know, they were there, you know, they like flipped the switch and we were able to keep going and like ingest all the data, which has been amazing.
36:54And then thinking about as we try to think about security and those building blocks there, we went through the SOC 2 process and everyone told us that it was going to take us six months to do SOC 2. But we were able to finish it in three months because of the fact that the core building blocks around security and keeping data secure were already available through AWS. And then even now, you know, in terms of like as we go to market, just AWS has been there to support us kind of at every step of the way. It's been an awesome, awesome partnership to have them in our corner. Awesome. So shout out to AWS.
37:24Shout out to AWS. Yes. Awesome. Well, once again, thanks so much, Cody. Awesome. Thanks, Sam. Thanks for having me. All right, everyone, that's our show for today. To learn more about today's guest or the topics mentioned in this interview, visit twimbleai.com. Of course, if you like what you hear on the podcast, please subscribe, rate, and review the show on your favorite podcatcher. Thanks so much for listening and catch you next time.
From the publisher
Today we’re joined by Cody Coleman, co-founder and CEO of Coactive AI. In our conversation with Cody, we discuss how Coactive has leveraged modern data, systems, and machine learning techniques to deliver its multimodal asset platform and visual search tools. Cody shares his expertise in the area of data-centric AI, and we dig into techniques like active learning and core set selection, and how they can drive greater efficiency throughout the machine learning lifecycle. We explore the various ways Coactive uses multimodal embeddings to enable their core visual search experience, and we cover the infrastructure optimizations they’ve implemented in order to scale their systems. We conclude with Cody’s advice for entrepreneurs and engineers building companies around generative AI technologies.
The complete show notes for this episode can be found at twimlai.com/go/660.




