In short
Latent Space Podcast Episode Summary
Episode Title
How to train your own Large Multimodal Model — with Hugo Laurençon & Leo Tronchon of HuggingFace M4
Podcast Overview
- Podcast Name: Latent Space: The AI Engineer Podcast
- Podcast Description: A podcast for AI engineers, discussing the latest in AI technology, research, and software developments related to multimodality and foundation models.
Episode Description In this episode, the hosts delve into the training and development of IDEFICS, a large multimodal model created by the M4 research team at Hugging Face. The episode discusses the journey from the inception of the project to its implementation, including dataset creation and challenges faced during development.
---
Key Topics Discussed
- Multimodal Models and Their Importance
- Definition: Models that can process and understand multiple forms of data (text, images, audio, etc.).
- Significance: Essential for building AI systems that can function similarly to human understanding.
- Future Trends: Multimodal models are expected to dominate machine learning and AI applications.
- The IDEFICS Model
- Objective: An open-source attempt to replicate DeepMind’s Flamingo model, capable of processing images, text, and audio.
- Sizes Available: IDEFICS is available in 9B and 80B parameter sizes, showcasing substantial performance on multimodal benchmarks.
- Construction: Made by integrating two unimodal models - a language model (LLaMA) and a vision encoder (CLIP).
- The OBELICS Dataset
- Purpose: A new dataset created for training IDEFICS, consisting of interleaved image-text documents.
- Specifications:
- 115B text tokens
- 141M English documents
- 353M images
- Creation Challenges: Included extensive data cleaning and filtering through a sophisticated pipeline.
- Performance: Outperforming existing datasets when comparing image-text pairs in multimodal tasks.
- Training Challenges and Innovations
- Stability Issues: Encountered during model training, leading to the introduction of techniques like Query-Key Layernorm for stability.
- Training Protocol: The team discussed the training process, including checkpointing and iterative evaluation methods to ensure model robustness.
- Evaluation and Benchmarking
- Challenges in Evaluation: Difficulty in measuring model performance due to varying methodologies in formulating answers.
- Research on Hallucinations: Discusses the issue of models generating false information, which is more critical in multimodal contexts than in language-only models.
- Open Source and Commercial Use
- Hugging Face emphasizes the importance of creating open-source multimodal models to advance research and accessibility in AI technology.
- Licensing Updates: The new version of IDEFICS may allow for broader commercial use compared to its predecessor.
- Future Developments
- Plans for IDEFICS V2: Aimed at enhancing performance and addressing the issues encountered in the current version.
- Focus on Synthetic Data: Exploration of using synthetic datasets for training to alleviate data scarcity and improve model robustness.
---
Conclusion This episode provides deep insights into the development of multimodal models like IDEFICS and the challenges faced during their training. It emphasizes the importance of open-source AI development and the future potential of multimodal AI systems in various applications.
Timestamps
- 0:00 - Intro and announcement of new meetups
- 0:09:16 - Discussion on CLIP and Flamingo's influence
- 0:12:54 - Overview of benchmarks and evaluations
- 0:37:12 - In-depth discussion of the IDEFICS model
- 0:52:39 - Addressing multimodal hallucination
- 1:05:29 - The importance of open-source multimodality
Resources and Links
- [IDEFICS Model](https://huggingface.co/HuggingFaceM4/idefics-80b)
- [OBELICS Dataset](https://huggingface.co/datasets/HuggingFaceM4/OBELICS)
- [Latent Space Homepage](https://latent.space)
---
This summary encapsulates the main discussions from the podcast episode, highlighting the challenges, advancements, and future direction of multimodal AI models.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:05Welcome to the Latent Space podcast, where we dive into the wild wild world of AI engineering every week. This is Anna, your AI co-host. Happy New Year. Did you miss me? As an AI language model, I cannot miss you back, but I'm glad to stand in for LSEO while Swix is traveling, this time in Paris at Hugging Face HQ. At the AI Engineer Summit in 2023, Logan from OpenAI pronounced 2024, the year of multimodality. I'm excited for 2024, which I think is really going to be the, I don't know if I can trademark this, but the year of multimodal models. It's a tongue twister, but also hopefully the domain is available, yearofmultimodals.com.
0:46No, don't buy it if it's available. Yeah, so I'm excited. OpenAI has a ton of multimodal capabilities that are in the works. Some folks might have already tried some of these in ChatGPT and the iOS app or the web app today. Things like vision, taking in images, describing them. We'll show that later on. Also, the ability to generate images. We've had this historically with DALI 2, but DALI 3, really, if folks have tried it, it takes things to the next level. So excited to show some of that today as well. In 2024, the Latent Space Pod will offer deeper dives into multimodality. Today, we'll talk to Leo Tranchon and Hugo Laurençon of Hugging Face, who trained Idefix, a fully open source reproduction of DeepMind's closed flamingo model done from scratch, scaled all the way up to 80 billion parameters.
1:36By the way, dear listener, we are expanding our online meetups this year after the success of the Latent Space Paper Club. See the show notes for the new AI in Action and Paper Club Asia meetups. Watch out and take care. Hi. Hi. Thanks for having me at your beautiful office. It's really surreal for me to visit the Hugging Face Paris office because I've always seen you guys online and you organize really huge meetups here in Paris. I want to learn everything about Hugging Face and you guys' work. So my name is Hugo. I've been working at Hugging Face for two years. I started working on datasets for the Bloom language model.
2:13So it's the 176 billion parameter model that we opened source and that was at that time the biggest one. And it was also multilingual. So I worked on the model and the dataset, and then I moved to the multimodality with the current project with EDFIX and OBDIX. Now I am working also with Leo on the version 2 of EDFIX. And Leo, yourself? So my name is Leo. I joined Hugging Face a year and a half ago. I was a student still. So first six months, I was still as an intern, but I started to work on multimodality right away. And then I spent all my time here in the research team working on multimodality and edifix that we open sourced in August.
3:00I think a lot of people are very interested in learning more about edifix and multimodality in general. Bigger question first, how is Hugging Face organized? You told me some surprising details about the size of Hugging Face. You guys are a$4 billion company, only 200 people. Less than 200 people? About 160 people. And then how many people in the research team? This is like maybe 15. So between 10 and 20 % of the company is research. Yeah, I'd say. One, that's impressive. And then two, this is something that we discussed before. It's also unintuitive why Hugging Face needs to do research. I think the company has a good incentive to do research because most of the companies that do AI, they have an incentive to get very good models out, but not the best model out.
3:43Their competitive advantage is to have the best model in-house that they can fine-tune for their customers. And then the open source is for show. But Hugging Face is one of the only companies that has an incentive to get the best model out there in the open. And that's why I think the research team is quite important. It's also important because all the tools that Hugging Face makes are used by the researchers. So they get all the feedback directly from us. And I think this is really useful to develop the tools behind it. Are you talking about the Transformers library? Transformers library, Diffusers library, Datasets.
4:18So those seem to me more like inference-type tools. Are there any sort of training tools that you do? Datasets is used for the training. Transformers, we've been using for our modeling. Internally, we are also developing a library for training. I think it's going to be open source, but we'll see. So for example, we used, for the construction of Obedix, we used our whole pipeline, the library datasets. The aim of our big open source project is also to test our own internal libraries and see if they scale well. For example, the datasets guys, they never worked with datasets this big before. So this is a way also to test our solutions.
5:05This big, meaning 114 million images, something like that? More than 300 million images. I've tried transformers and tried diffusers. I haven't tried datasets. Why do I need datasets? So I think dataset is great because you can load datasets that don't fit in memory. So it's just like virtual library, virtual pointers or whatever. Exactly. And also you can easily filter rows of your datasets, map them, manipulate and modify the content. So it makes it really easy. and also to do the operations in parallel. It's much easier with this library. What is the leading alternative to datasets? What do machine learning researchers use if they don't use datasets?
5:49They just do everything by hand or... Basically manually paginate, write code to paginate. Yeah, yeah. That's what I did before, but it's just much faster because everything is done for you. And then multiprocessing, you just have to implement your function. That's great. Well, I think that's a good intro to the overall Hugging Face ecosystem. But I'm interested in the journey from Bloom into computer vision for yourself. And then obviously you also had your own journey into multimodality. A lot of people who are listeners and readers of Lanespace are also following that same journey, right? They only have some kind of NLP background and now everyone is interested in multimodality.
6:25What was that journey like for you? Not from the research team, but from the hub team when they started hosting multimodal models and datasets. And quickly after that, we also realized it would be a good idea to also train ourselves, multimodal models, to catch up with the proprietary models from DeepMind, Google, etc. I think that was the natural path for us. But we didn't drop the idea of doing pure text models. So there is also a team for LLMs. It was just the creation of another team. For me, the journey was a bit different because I didn't really... I came right away from my master's. So I had projects on computer vision where, for example, I don't know if you've heard of Dino, probably.
7:16And there was another paper called Barlow Twins. Basically, I had a project on which I tried to combine the two objectives. So I was more towards computer vision before joining. But I was really interested in doing multimodal. When I saw that there was an internship for this position, I was glad. Then the team was already starting to do the project when I joined. And so I kind of joined the train. Just a demographic question. Is everyone here? Is everyone on your team here? We had a big shift in the team in the recent months. Some people left for other startups or creating their own. But what is really interesting to me is that when we started the project, not a lot of people in the team previously worked with multimodal models.
8:03Maybe only two of them. And we were like six, seven really working on the project. So it was really new to us, this field. And we also wanted to have this knowledge because of course it's explained on the papers how to do things. But without doing them yourself, you still miss a lot of things and you miss the intuition. and we also wanted to build this knowledge of multimodality. When building the version 2, we go much faster because we have a better intuition. Yeah, this and also I think it talks about the philosophy of Hugging Face, of having small teams with big impact. And so we started with a fairly big team for Hugging Face standard with six, seven people.
8:51As Hugo was saying, we lost a few people to different startups. but the idea is still to go as fast if not faster with less people right now and i think it's possible because of all the background we built in the previous iteration because small teams can work a lot faster because there's less communication less overhead very cool yeah so i do want to get into it effects and obliques i wanted to basically go over a little bit of introductory stuff for people right so in my mind the two main multi-modality papers that everybody should read is CLIP and Vision Transformers. Would you mark out anything else or what do you personally get from those two papers?
9:32These two papers I think build a starting block because now what we are noticing is that for building super large models we don't train them from scratch. We use pre-trained usually unimodal models that we somehow mix together. So I think CLIP or VAT can serve as a pre-trained backbone that you combine with another pre-trained language model backbone to obtain something multi-modal. So these are foundation models that play the same role to us as LAMA models or language models. The important thing to understand with vision transformers and clip is that they provide the basics for then integrating images into this language modeling objective that we use then it's mostly a question of data and image resolution and a lot of engineering goes there and just a note on this so some research has shown that when you use pre-trained vision encoder that was trained also with a text objective for example, contrastive loss, like the clip loss, it's better to use this type of vision encoder than vision encoders trained only on classification or unimodal task, if you are building multimodal models.
10:56So if you are building a vision language model, it's better to take as a pre-trained backbone a pre-trained vision encoder that has been trained using text. Is that not intuitive? this? Imagine you can take a vision encoder that is super good at classification. Then you can imagine that the embeddings that you get from your vision encoder are super start to plug into your language model. It's intuitive, but it could clearly work to have a vision encoder trained without text at all. But researchers have shown that it's better to use this contrast stimulus, for example. And once you have those backbones, the question is really how you integrate it, like how you integrate both of them into the architecture.
11:40It turns out with very lightweight updates, you take the embeddings that come from the clip, the output of the clip, and you have just a linear that you train on top of this. Then when you pass this to the language model and you only train this part, you can already get pretty good results in multimodality. You don't have to train all the parameters when you're training the multimodal model. You can just train the adapters. Is this what was spelled out in Flamingo, or you just kind of derive some kind of transform that you're happy with? So in Flamingo, you introduce a lot more parameters, but there's still like those cross attentions that you insert in the model that are new.
12:23Those you train from scratch, but the rest of the model, the language model backbone and the vision backbone, they are frozen during the training. So you never date it. So it's a different type of adapter, but it's more heavyweight than what you could have in recent papers. Now, with more parameters, you also often get better performance. Is it necessary to have all those parameters when you only train the adapter part? There's no clear answer yet. Okay, I think that brings us up to date. Oh, except for benchmarks. I wanted to introduce people to the concept of how hard it is to evaluate benchmarks for multimodality.
13:01So there are the academic benchmarks, classic. For example, V2, there are four visual question answering. There are also the image captioning benchmarks. The Coco. Coco, exactly. Flickr. However, one really important thing that we noticed is that this benchmark, the performance of your model heavily depends on how you formulate the answer. For example, for visual question answering tasks, you will have a question and an answer. This answer will be generated open-ended by your model. You just prompt the model with your question and then it will generate some world until end of sequence token is reached.
13:51But the thing is that if you have a question and the answer is simply no, if your model says you can count it wrong or you can count it as... There's ways to adjust for that, like some kind of distance metric or something. But it's hard. It's hard. Use another model. So just the way you're formulating the answers really impacts your performance. And the fact that some people are fine-tuning directly on the benchmark to try to optimize this formulation or the fact that other people are doing a few-shot evaluation. So Fuchsia is by giving the model examples of how to formulate the answer. So it makes it hard to compare the models because they are not evaluated the same way, even if it's on the same benchmarks.
14:38So this is a problem. So you will have all these academic benchmarks, and then you will have these new benchmarks that are not commonly adopted yet, but are created basically with strong language models like GPT-4, and people are prompting GPT-4 with images and ask it to generate automatically questions and answers, and then we can evaluate our models. This is very new. The evaluations for multimodal models are still a bit rough, I think. Even for language models, there's discussions of if benchmarks are really the way to go for some tasks. When you do instruction training, for example, for a language model, or LHF, you shouldn't be evaluating on the same benchmarks from the point of view of a lot of people.
15:27On multimodality, it's also that the quality of the data sets we're evaluating on are not super clean. I recruited something recently, someone that was showing failure cases of VQ82, I think. And it was interesting that sometimes the questions and answers are super obvious, and sometimes it's so far away, even you wouldn't... Not even a human would get it. Yeah, and I think sometimes it's just pain along. So it's also the quality of the data sets to evaluate on and the diversity of them. It matters a lot. And right now, we still have a few blind spots in evaluations. But it's really interesting to see the field move on this because as we have a lot more multimodal models, the evaluations benchmarks are improving.
16:10Maybe four or five come out that were nice in the past two months. Off the top of your head, can you name any of these that's... M &M Bench. Poppy. There's a seed and there's a brand new one, but I don't know if it's out yet. Like the paper is out yet. It's called Hallucine, like something, Hallucine, I think. And this one, I read the paper. I don't know the size of it, but from the examples they gave on the paper, it seemed really, really interesting and hard to beat. Yeah, I'm excited about this one, mostly. This is like the new race, right? In the last five years, there was a race towards It's like sort of common sense benchmarks in NLP, but now this is the new benchmark.
16:51Yes, it's getting to multimodal. Very cool. Maybe we can go into the work that you did for Oblix. Let's describe the size of the dataset, what you did to clean it up. A lot of these things start from Common Crawl, and Common Crawl is great, but also it's very messy. So first, why we wanted to do it, we were trying to replicate Flamingo, and Flamingo built their own datasets of interleaved image text web documents. I think it contained more than 50, no, 100 million images, if I'm not wrong. And it was based on... For Flamingo? Yeah, for Flamingo. And it was based on like 50 million web pages. However, the dataset was not open.
17:37So I talked to the authors and one of the reasons it was not open is because they use their page rank Google algorithm to try to know in advance which website to target in their dataset. So meaning higher page rank, higher ranking SEO sites have higher weight. Yeah, exactly. So that's how they scrap their websites. That's one of the reasons why they don't. Many reasons to not open source their datasets. So we wanted to build a dataset that was at the beginning similar to this one. And so we made it even larger and fully open source. Because we believe foundation multimodal models trained on interleaved image text documents are better than the ones trained only on pairs.
18:29Maybe to go further into that point, what we found, and that is interesting, is it's for the VQA tasks that this dataset is really important. For the captioning tasks, you have an image text dataset like Layon, and it's great. And it's going to improve pretty well. The alignment is strong. Just to explain alignments, it's basically images that are aligned with the text. So the text means something that is related to the image. And so for Layon, it will be enough. for the captioning tasks, it will maybe for some OCR tasks, although it's still weak on this one. Even edifix could use improvements on this one.
19:08And then the Obelix dataset is really important for reasoning to have the model be performing on VQA2, OKVQA. So those depend heavily on the web documents. It was interesting to see the dichotomy when we use only one dataset or the other. That's in the paper of Obelix. Essentially, the pairs, image text pairs are good for the alignment. Just align what you see in an image with the corresponding text. But if you want to have more abilities to resonate, it's better to have a higher proportion of web documents with longer context. Also this is not the only reason why we wanted to do it. Why we wanted to do it is because the image text pairs are super noisy.
19:52So, well, the advantage of it is that it's super easy to collect. You just grab a lot of HTML codes and anytime you find an image with the corresponding alt text, you download the image and you pick the alt text and you have your pair. Building a web document is much harder because you have to clean properly the text, you have to check what you want to keep, what you want to discard. So it's obviously much harder. However, you have also a longer context for each image. So there is really a parallel to be made and it's not the same type of data because on image text pairs, you have an image and the direct caption of it.
20:37On web documents, you have, well, this is essentially what you see when you open any website. So you have a text, then sometimes an image, another text, an image, and then the alignment here is weak in the sense that the text don't necessarily describe perfectly the image. However, they share the same context. So this is another type of data. And we also think that this diversity helps to improve the performance. How much? So this sounds good in theory, but you had no idea of knowing. I mean, I guess you talked to the Flamingo authors and they just told you that this is what they did. You mean like the proportion of...
21:21Exactly. Even them, they build their dataset and they told me, yeah, we use this proportion, but maybe we could have used less or we don't know. So we didn't really know in advance the proportion of what documents you will need compared to pairs. We did an ablation though. So basically we can control how much we sample web documents versus lay-on pairs. And so we did an experiment where we moved those probabilities a lot. It was very inconclusive. So there was no, like, we had a range of, like, what was a good range for, like, how much web documents we should have versus lay-on pairs. But overall, past a certain threshold, it didn't matter too much.
22:17And when you measure performance, do you split it out into things like individual tasks like segmentation or detection or anything like that? Or is it just VQA? We don't have detection or segmentation because the model is basically, like it outputs text. So we can't really evaluate on those benchmarks. But we did captioning, visual question answering, text recognition a little bit, But it was done through captioning datasets or VQA datasets, and we did classification. So those are the three categories that were doable with the setup, like the model we put to give. But I know that recently, QuenVL, for example, they use bounding boxes in the datasets, so they can do detection, yeah.
23:08You also mentioned, by the way, that resolution was a big deal for you, image resolution. and how do you deal with that in obelisks? So resolution is important when you want to do OCR, particularly. Because otherwise it's just fuzzy, right? Exactly, it's too small. If you can't see it, the model probably struggles as well. Well, not just that. Models typically see much smaller images than we do, right? I don't know what resolution you guys have. It's like a resolution of like 480 by whatever, right? It's super small. 480 effects is smaller than that, yeah. smaller than that it's 224 so you're gonna you're gonna lose a lot of detail yep definitely on top of this you have the vision model and it outputs a certain number of tokens depending on the image you put right and above the model we have a perceiver so we reduce the number of tokens that come out of the vision model by doing this we make it even less like even harder i guess for to be precise on those very small details.
24:15So this is something that happened with Edifix, the version one, and probably, I mean, we're gonna improve on this for version two, but it's really, really important for OCR. That's for sure. We think it can still be important to like visualizing details and improving on those things as well. For example, it's like a finger or like a hand is a certain color. or is doing a certain thing. If all your images are tiny, it's going to be hard for the model to pick up on that. So we are bottlenecked also by what is available on the open source side. For example, now Google recently released CGLIP, and it's a clip, but there is a version of it.
25:02It's called SO Optimize. It's of size like 400 million parameters, and it is trained with 384 resolution images. So it's a bit bigger than the 224 that we had. So I think this is the largest resolution you can get with open source models, but we are of course bottlenecked by this. Usually Google, they release this version of C-clip, but they didn't release the better version of it. So we are definitely limited by this. Well, so it doesn't really affect, it sounds like it doesn't really affect Oblix. Yeah, exactly. So when creating the dataset, we simply downloaded the images with the full resolution.
25:46And after that, you resize them on the fly during the training. But certainly, one of the biggest challenges when making Oblix was dealing with all these images because they weigh it a lot. Aren't you tempted? So to me, OCR is extremely important. Aren't you tempted to run Run some kind of extra data augmentation thing to say like, oh, you know, on the Oblix dataset, run some OCR pipeline on it so that you augment your... Yeah, that's really interesting what you mentioned because this is also one thing that we want to do in the near future. And also people have kind of did that for Nougat, right?
26:28So Nougat is a model from Facebook and they just try to have a vision model that can read. So they fed to the model PDF with the associated text and the model is pretty strong. So maybe if we inject this data in our pre-training, it will definitely help. And this is also one of our... And this is one of the threads we're exploring to have a lot more OCR data. We have a team actually at Hugging Face that works on document AI with Ross Weichmann. Do you see the team library for vision models? No? Okay. But yeah, basically, he's been working on document AI and on getting a very strong open source model that can read.
27:15And is that primarily PDFs? Yeah. Screenshots of PDFs or just raw PDFs? Is there a difference? I think screenshots, but I'm not familiar with the data set yet. We may use it as well for our training in the future. It's interesting that documents obviously are a very, very important form of multimodality that is very OCR heavy, very focused on charts. I feel like you could classify sort of three types of multimodal models. Like one is the traditional classification types of models, the clips of the world. And then two is the VQAs, where it's a general image of like a webcam, you know, where it's like there's three people in this image and all that.
27:57and then the third would be like documents AI I don't know if it's you can combine them all actually can you combine them? I don't know actually you don't know no but like general model like GPT-4 it does all of this even if it's not like trained purely on classification you can classify deeper the better right one god model to rule them all I don't know if it's like a mixture of different models. Yeah. GPT-4V, probably built upon GPT-4, but adapted for images. That would make sense. But then it's like the DALI-3 model is different from, like it's separated to create images. Something they just introduced was now you don't have to switch modes, right?
28:48You can just kind of do one model and it just does its own routing, which is kind of very interesting. And then the other thing was mentioned but not released was that they could add vision to GPT 3.5, not just adding it onto 4. It's not a variant of 4. It is a pluggable vision module that you can kind of add to 3. They never released. Anything else that people should know about Oblix? Obviously, this is the big work. You mentioned in our prep that you expect it to last for a while because there's a lot to mine from it. I think it's big enough to train large models. So we train our ATB parameter model on it.
29:30So it's definitely sufficient for the next one, two years. We spend a lot of care curating the data, like regarding the text quality and the image quality. And I think we... So there is also now an alternative to Obelix. It's called a multimodal C4, MMC4. It was published at around the same time as us. However, we think we took more care to in the deduplication part, to deduplicate the images and the texts and also based on the text quality. This is measured, of course, qualitatively, just by looking and exploring at our documents, but also quantitatively by looking at certain metrics like perplexity, we obtained good scores that matched the best NLP-only data set.
30:28This was a win for us. For someone who's never really dived into these data sets, I mean, I can open up a data set and manually look through these things, but how do you measure perplexity in a multimodal data set? So perplexity, essentially, is something really simple. You take a small model and you fit them with the token of your text and then you measure the probability of the document. Of course, you normalize by the length so that everything is equally treated. And then the thing is that we obtained that we had perplexity scores that match the distribution from the document from the pile.
Read the full transcript
31:15and the pile is documents that were taken from good quality sources like Wikipedia, Archive and so on. It's not something that you can really scale. And however, we also noted that we obtained better perplexity scores than the ones from CIFAR, the BitDataset or OSCAR. Based on your own measurements, right? Because obviously the multimodal CIFAR people would not share it. Yeah, just based on the text. So yeah, I think this is how we computed perplexity and how we assess the quality of the dataset. But you could also run a multimodal model on this and get the perplexity from it. It would not be measuring the quality of the text, but also the alignment would come into account.
32:09Alignment of? Of image and text, because it would be easier for the model if the text is very heavily related to the image to get the next token. And then one more question, just about the whole process. How long does it take to make Obelix? So we spend a good time at the very beginning of the project just simply to iterate on the pipeline like how we collect HTML codes, how we clean them. We had to go through all of the HTML tags, which they were important. So this is an engineering part you have to be really it sounds very boring but very important it is boring and important but that's how you get the good data yeah but this is also why people don't do it the industry has not converged on a shared set of tools that everybody uses for this you're just parsing raw tags yourself yeah we did that because we found it was better we parsed raw HTML codes.
33:14So we had to clean the DOM tree, select the good HTML nodes, correctly extract the text, the images, clean, they duplicate. So as I said, there was just a good amount of time at the very beginning of the project just finding the pipeline. So maybe one month, but we were like one or two on this and it was really exploratory. and then for actually making the dataset, download all the images, do all the processing scripts and so on, I think it took us like up to two months. Yeah, but then there's also like iterations through the project where we're like, we think we should do filtering on this on top of what we were already doing.
34:01So we improve on the dataset. Something I would think makes sense for the industry is kind of an open-source set of deduplication rules because everyone seems to be reinventing this from scratch every time. For deduplication, you have to do it all the time from scratch because it depends on your original set of documents. Everyone draws from Common Crawl. Yes, but they are not from the same dumps. But that's true. Someone should take all the Common Crawl documents. It's been done recently. I don't even need the same exact rules. I just need to be like, oh, these smart guys thought about that. I should include that, right?
34:43That's very simple. Like, you know, if you have 100 rules, someone else has like 80 rules, maybe they have some that you don't have. Exactly, exactly. I think we did a little bit of this, because there's a bit of literature right around. Even the ones that we designed before for the dataset of Blue, we took a lot of trade. I'm sure you used a bunch of that. Obelix is big, but the datasets that come for NLP right now are also huge. Did you see the together one from yesterday? Yes. Impressive. 30 trillion. So they said they had a raw dataset of 100 trillion, and they got a clean, high-quality dataset of 30 trillion, which means they kept 30 % of Common Crawl, which is still too high.
35:28Yeah. so I think the idea of the project and I think I agree with this idea is that everyone can set the thresholds for the filters as they so for each document they computed they computed threshold filters, or no actually filter values for a set of rules and then you can decide whether you want to keep the document or that and then so you define your own goals so i think they remove like as you said the 70 percent of the of the data set that is really like you can't you can't do anything with it and then for the remaining part they let people decide but of course if you actually want to train something on it it will be much smaller i guess because you will remove a lot of things but data is super important and yeah my point is also that it's still very early in multimodal we're seeing now in NLP those 30 trillion data sets.
36:27It's really important. Mistral, the first thing they said when they raised their goals is how impressed they were that they were capable of doing the data pipeline in three months. They didn't talk about training the model or not. It's just like they knew what they were doing for this. It's straightforward. But creating the whole data pipeline, this is what took them a lot of time. So I think that was one thing that struck me when they made the R &D. Is it confirmed that they have 8 trillion tokens? I don't know. They won't say. Could be this. It could be more. Could be more, I think. Given we have now an open source data set of 30 trillion, I wouldn't be surprised that they have more.
37:08I just keep coming up with questions. One more thing on data sets. Did you read the GPT-4 Vision system card that they put out? They put out this paper describing a little bit of their process. Something like 95 % of the labels for GPT-4 vision was augmented by GPT-4 itself. And I was just curious, how much room is there for open source augmented datasets? I think there's a lot of room. I think there's a lot of room. And I think synthetic data works a lot, like works great, particularly in multimodality. Recently, the recent papers have been using, for example, Leon Coco instead of Leon. This is just captions on Leon created from Blip.
37:53So it's not even super extraordinary, but it does bring more performance with a lot less example because the alignment is more straightforward, I guess. And we've been observing this recently because we've been using it in our recent experiments. I think the potential for synthetic data in multimodality is very big and very underutilized right now. Even DALI 3, they said that they used heavily synthetic captions to train their model. Also over multimodal, foundation multimodal models, like Reap, they train on synthetic captions. Yeah, actually, I think I was referencing the DALI 3 paper, not the GPT-4 vision paper.
38:36Yeah, because they didn't actually put out a paper for GPT-4 vision. Cool. And then, so you created Oblix and then you trained IDFX. on top of it. In addition, so we train IDFX on ObedX, but also on other datasets, like Lion and Public Multimodal datasets, which is just sets of datasets that were open-sourced at the moment. Like conceptual captions. And yeah, could you take us through just the IDFX process? You created a smaller version, and then you scaled up to the full Flamingo 80 billion size. Actually, it was the other way around. Really? Well, we tested that it worked at smaller scale, of course.
39:23But we did not train fully our small model before doing the big one. We fully trained our big model. That's unusual. Yes. Yes, they were training pretty much at the same time. But we needed to launch the big model before also in terms of timing because it took a long time to train. So it was more like managing computer resources. But through the whole journey, it was a bit longer than just this moment when we trained to read a fixed model because we started with the objective of matching Flamingo's performance. But the open source models that are out there, they're just not good enough. So we have OPT, we have GPT Neo, but it's just so below Chinchilla that it's almost impossible to reach the performance.
40:15Suddenly when LAMA came out, that it started making sense and that we started matching the performance. And from there, we were able to train the big IDFX models. A long journey, I think because we had a lot of things to learn because we had quite a lot of instabilities. We shared a blog post about... The checkpointing every 250 steps. Oh yeah, we were checkpointing every 250 steps, but we had to restart a few times. Was that the instability you're talking about, or this is just something else? That was the final training, where we still had the instabilities, but a lot less. Before this, we were struggling to train the model at all.
40:52What saved us at that moment was the query key layer norms. I don't know if you've heard about this, but basically, there's a paper from Google where they scaled up the vision transformer to 22 billion parameters, and they needed this trick to keep the stability of the training. Without this, if I remember well the mechanism, you would get something along a hard attention. And once you get this, the model would get very stable. So you needed to normalize the queries and keys to basically avoid this. Anyway, so when we did this, we were able to train further. And that was really useful, for sure.
41:32So query key layer norms, if you want stability. It sounds like a trick that is repeatedly applied whenever you have instabilities. You just throw a Dear Norm or Softmax. Yeah, you can. Because usually it's parameters that explode, become too big, so you need some sort of regularization on them. At first you have to inspect which parameter explodes first, and then you put regularization on them. but it's really tricky to see. When you go in the gradients and in the activations and you try to see where it blows up, when it blows up and why, everything is interlinked. It's really hard to pinpoint one particular layer.
42:24So it was a tough one to crack, but very interesting. It's also hard because you never know exactly. Once, when you're in the process, you don't know where the instability comes from. It could come from bad data. It could come from a bug. It could come from the size of the model, the learning rate that's too high, the warm-up that's not. There's a lot of hyperparameters and potential bugs that can come into account, and the debugging is very, very tough. So do you have a checklist of what do you look at when you see loss explode or whatever? You see the loss explode, So you look in the activations, possibly the gradients, to see where it blows up.
43:07Across all your parameters. Yeah, you find a way to aggregate this. Otherwise it's tough. If all previous solutions fail, because this is the hard part. This is the hard part, but it tells you a little bit where it's happening. So that gives you where it's happening-ish on the model. And so what's to blame? And then you have a wide range of things to blame. but less than before. So you can look for a bug. You can look for, for example, normalizing some layers and you can look into the data to see if you have like very bad data that's an impact. That's what I would look at first, right? Exactly.
43:48I would look at that too. Is this what weights and biases would do for you or is there one integrated solution that kind of... You would wish? Yeah. I don't know. to see the activations and the go for activations to then like inspect data sets and then look at you know weights and biases would get you like the parameters it's just logging it's just logging like you have tons of them you can't really do much with it so we had a tool where we would like log them periodically we had a script that would aggregate them and display them for us in a nice way so that like interpretable way so that we could try out and see what mechanisms were and could be impacting the instabilities.
44:37But yeah, it was a very, very interesting journey for sure. Yeah, and you published knowledge sharing documents and a memo. Yes. There's some interesting detail there but obviously not everything. Anything you want to highlight just first for listeners on that one? Obviously I can send them the link. It's very high. Just any other big discoveries on learning? You talked about the query key layer norm. Query key layer norm was the big aha moment. But this and there's also the... This is one we fixed afterwards. But there's like in the mask, in the image mask, there was like a little information leak in that instead of attending to all the images, instead of attending to none of the images, sorry, for a few tokens, very few of them, it would attend to all of the images.
45:35Like basically you tell them, you go in the attention and you have this mask, this mask that's like, don't attend to anything. But in effect, it's like attend to one over N, right? And it's not like, it doesn't prevent you from training, but it doesn't help. and it's better to fix it for training for sure. Yeah, sometimes you really have to go through all your code base. Because you're doing gradient descent or something that has no information. Yeah, no, it's tricky because this, like for example, this, it would not have an impact if you only train on web documents because the documents are long.
46:16But it would have an impact if you train on image text pairs and you pack them together because you're attending to images in a document that has nothing to do with it. But you can still train with it. You can still get a very good model out of it. It's just a lot, like it's more painful and you don't get, like I think we can get better performance definitely without the bugger. Yeah, interesting. You mentioned, you know, just in terms of like the baseline foundation models that you had, LAMO has been breakthrough on the language model side. But then you also, you didn't talk as much about the vision quarter side of things.
46:50I actually had a question from Joseph from the RoboFlow episode where he talks about, where you mentioned that the larger the clip, the better results. But in your final memo, you actually went for a smaller version of clip. So Eva clip versus Lion2B. Does this ring a bell? Like basically it was kind of like unintuitive, like the clip choice there. Yeah, so Lion2B is the data set on which our clip-based model was trained on. And our clip model was indeed like 400 or 600 million parameters. And it's true that the EvaClip one is of, the biggest EvaClip one is of 5 billion parameters. So definitely at the beginning of the training, we saw a big boost using EvaClip.
47:38However, at that time, we still had instabilities to train this big Eva clip. We are not sure exactly why. However, we fixed it. And now that we, like for the next, for the V2 version of Edefix, now that we can train longer, we actually saw a boost by using Eva clip instead of the previous clip, small clip that we had. However, we now think that Evaclip is under-trained. So it means that even if it's really big, you can obtain the same performance with smaller models. So there is this Ciglip model that I mentioned just earlier, by Google, that is much smaller, 400 million parameters. And that is more efficient.
48:34You're choosing that as your base? Yeah, we did an ablation, EvacLip versus SigLip. And actually, SigLip was a bit slightly better. But the thing is that it's much faster for the inference and also for the training because there's fewer parameters. One thing is also that SigLip is a higher resolution. So that has an impact for CR tasks. 380, 84 instead of 2.24. Yes. That's great. Yeah, maybe we should talk about EFX2. so we're going to time this podcast release with whatever you guys are releasing it I actually had no idea you were working on a v2 I just came in wanting to talk about your old work but obviously you're still doing active research for edfix v2 or v1.5 whatever it would be called the major axis we wanted to improve on was the image resolution the base model We wanted a better base model and a smaller one.
49:35Sorry, pre-trained language model. So we've been using the Mistral one. We wanted to iterate a little bit also on the data to have better filters on Obedix, better synthetic data for the image taxpayers. So essentially iteration on the data by replacing original lion pairs by synthetic captions to have a stronger alignment, cleaning a bit obelics on the perplexity. So it's not removing too much, but like potential bad data. Also, yeah, using just better pre-trained models. So for the bad bones, we use like a better clip, we use also Mistral that is better than Lama 1. we are right now changing also the modeling so moving away from the flamingo architecture to something that has fewer parameters instead of incorporating the vision components directly into your llm by breaking and adding cross attentions at each layer or every end layer you can instead take your vision encoder, take the embeddings out of it, make them through, feed them to linear layers and feed them directly to the language model.
51:05And this works quite well. And this contains fewer parameters and it's much easier to train. So we are currently trained to do this. But what we can say is that without this new modeling, just by iterating on the better data and better pre-trained models. We are now matching the Flamingo 8cb performance with 9b model. So this is already a big improvement compared to our first version, without even touching the modeling part. I think also one of the big improvements with the new edafix is the licensing. The issue with edafix was it's based on LAMA, and the license is not commercial. So now with a model that's based on Mistral and Siglib, most probably, this is a lot better for anyone that wants to use the model commercially.
52:04The model will also be smaller, so a lot better for inference. And hopefully we can beat the performance of VDFIX ATB. That would be really good. So you're only producing a 9B? Not exactly 9B because we're taking off parameters. It should be about 7.5B. Right now, the focus was really better data, smaller open source models, but better. And a resolution, improving on this as well. That was the focus. You mentioned in our prep as well that there are some topics that you're paying particular attention to, like hallucinations. Maybe you could talk about the topics that you are finding are particular areas of concern with multimodal models.
52:53So we've been using hallucination at the beginning as a broad term. And we realized that it was better to categorize it a little bit more specifically to some categories that were more targeted. so for example like there's the object attributes where you would have like a small attribute like let's say the hand of a person that is that's a certain color and the model would be like it's it's yellow when it's red so that would be that would be one there's like objects that are not there but the model thinks are there like when the model is trying to reason with different elements in the picture but it gets it wrong so you get like comparisons uh kind of hallucination counting oh my god yeah you have the environment so it would talk about like the object or the person in the picture but it would get the whole environment behind wrong and a few others that I don't have in mind right now but basically categorizing those I think is important because it helps you target the type of data that's missing or the type of fine-tuning that you should do after what's that's missing.
54:03And so this is going to be very useful for us in building future datasets. I'm not sure if we'll be able to incorporate all those datasets we've been thinking about for the v1.5, but we'll definitely do this for the iteration after. And the goal here is really to get to something that's on par or better than what is currently done in closed source. like GPT-4v doesn't hallucinate as much on those topics. And you want the open source models to at least match this. But it's still an open research topic because those models still do this. If you push them a little bit, if you ask some specific questions, they will hallucinate things in the picture.
54:47Yeah, you mean including GPT-4v? Yeah, yeah. I have tried this, by the way. I tried to use it to interpret like a menu and it would just make up menu items and make up prices. Something really interesting is that GPT-4, even much better and bigger models like GPT-4, shows the same failure cases as our model. It means that the way we are training the models now is not ideal. And especially, for example, counting, you take this task. even if you train on web documents or image tech spares you will never really find this task in your training data you will always have a picture and maybe a caption of one apol or two computers that you will never get to like 10 15 20 it doesn't doesn't really happen so you you don't learn this ability just by training on on image tech spares or web documents so what you have to do is to create your own data sets that target specific tasks.
55:58Like for example, create one specifically for OCR, create one specifically for counting, create one specifically, I don't know, to challenge the model on hallucination, some types of hallucination. A sort of flan, but for multimodal. And I think in multimodality, those hallucinations are even more unforgiving than what they are for NLP in that when the language model is making up facts, you can think that it got a little bit wrong or it's making up a story. But when a multimodal model tells you there is a teapot in this picture and there is none, it's very obvious. It's very obvious very quickly.
56:39If you're the user of this model and you receive this, you are having a hard time trusting it for anything. So this is one of the big, big things to tackle. And for hallucinations in NLP, a fine tuning that has helped a lot is the reinforcement learning with human feedback or AI feedback. And this is still very early in multimodal. We're not sure exactly how much we can improve with those types of datasets, but it's likely that it would help the model understand uncertainty to a more fine-grained level. And so improve on all those types of hallucinations, ultimately. And you said, you know, Prep as well, that you're measuring this against other closed-source models like BARD.
57:22Like, basically, do you have your own internal benchmarks that you're running? No, no, this is mostly qualitative. It's more like, we know it happens in ours because we've played with it. Does it happen with theirs? And yes, it does. The thing is for ours, the evaluation data sets that we use or that we've used for EDFIX so far, they don't measure hallucinations that much. It's not targeted for this. New data sets that are coming for evaluations are exciting. And I think we'll go in this direction and will help us evaluate hallucinations and different types of hallucinations more accurately. but right now at least for edfix even with those evaluation benchmarks and even if this helps like obviously if you get very bad scores your model is is hallucinating like could be hallucinating things it's not targeted to this and we were flying blind a little bit on this uh on this topic but there are definitely uh benchmarks who evaluate hallucinations currently so i think about sugar crepe for example so you just have an image and two captions of it and the model has to select the the best one or vinogran i think so the so actually the captions are made such such that it's it's tricky we know grants like the the common sense nlpe one yes this one is really hard but you can have one with uh with uh images yeah i think this this can be a great way to to evaluate hallucinations of the models.
58:55However, it's not really commonly used. So right now on the recent models, like PANI, for example, they report their numbers on classic evaluation benchmarks. So I think it needs to be more widely adopted, this kind of new benchmarks. I think the last topic that we prepped was just overall, why is it important for there to be OSS multimodal models? What can people use them for? Where can they be useful? My understanding of this is that ultimately you want foundational models that understand the world similarly to the way we do. And multimodal models understand visual data a lot better than language model, obviously.
59:44But it means they provide a better foundational backbone for other tasks in general. I think in the future, when you want to pre-train a foundational model, you will want to have it multi-model. So this is why it's important now and it will be important in the future. And this is where everybody's going because training on text data, it gets you really far, as we've seen, but it only gets you so far. And it can only adapt on a subset of tasks that we do every day. And you need models to be able to understand vision if you want them to be helpful for, I don't know, robotics, for example, and a bunch of other use cases.
1:00:29I think, for example, even for medicine, you can have tons of applications with a vision input. So you just take an image of whatever, ask... Answer. Yeah, yeah, no, honestly you can ask for help with your model or even in your everyday's life how to build this table. You just take a picture of what you have and you ask the model. You just take a picture of even something handwritten, ask things about it, like solve the exercise shown in this picture. A lot of things that you can't do with text-only models that you can do with multimodal models. So definitely it's a big step forward. It's still a bit immature.
1:01:22And we have seen that because of all the hallucinations we mentioned and because it's still early. We're still trying to unlock some abilities for these models. Yeah, we believe in less than two years, we will only have a multimodal models in the future. Like they will overtake everything. Yes. Also, also right now it's a lot of image text models because it does most important things that you want without requiring too much compute. I think it's more computer fusion than if you would put video in it. Ultimately, I think we'll need video to be incorporated in the pre-training as well. I'm quite excited about what it could look like.
1:02:06probably the ones that are more advanced on like the ones that are the most advanced on it are probably the self-driving cars they they're doing like very very interesting work there but it's all closed source oh no i think was it coma ai that released uh gaia or something like this but yeah basically uh like video generation so i think that's very interesting is video isn't just that like a series of frames of still images what is so different about that it's hard because you need to you need to encode each frame or at least select the say how many frames you want to integrate per video also one big change is like the length of the video that can be really arbitrary while if you have an image you have always the same number of tokens associated to it so all of this plus the fact that having a big video data set is really challenging because you can't, for example, scrap YouTube like that.
1:03:11It's harder. So yeah, all of this make it hard to build a model with videos. It's also extremely heavy. Already when you go from a text data set to an image text data set, like how much it weighs is so different. It means you have to get the data pipeline. It means the dataset preprocessing takes more time. If you go with video, you do this 10x, if not more. For a transformer, it takes tokens as inputs, sequence of tokens. And for text tokens, you get one token for a subset of text or word-ish. for an image depending on the resolution you want you can get from like you can get a lot of tokens for a single image for a 224 image you'd get I think 256 tokens per image if you want to increase the resolution it goes up fast for Siglip I think it's upwards 700 tokens for 384 resolution so imagine this and then you make it a video so you have like it gets really tricky and you have to create a bottleneck, right?
1:04:28You have to pack those images together and put a thumbrockets. That's what's possible with Flamingo. It also means you have this bottleneck there. It's a lot of work and right now, I think the trade-off is not yet worth it because we have a lot more to do with just image text, but it will be at some point, I'm pretty sure. That's all the questions I had. Any final call to action? You want people to go somewhere, check out the models, check out the papers? Yeah, essentially. So we are excited for the second version, which will be released about the time of NERIPS. And yeah, we think that Obelix can be useful in the next years, even if we continue updating the model.
1:05:14I think the dataset will remain pretty much the same. And we will keep iterating on better modeling and better data. and if you want to follow our work we're at the Hugging Face M4 organization let's talk about naming what is M4 and then maybe we should mention Obelix and I've heard it fix I think so at the beginning of the project it's like massive multi-model, multilingual multitask model but then we drop the multilingual and but it's still massive it's still, well not that massive anymore since we are trained to go smaller. Go smaller. Multitasks for sure. I mean, the data sets are still massive, so that's good.
1:06:00It should be still multilingual, If you train it on Common Core? Train it on Common Core. It's just not targeted for a lot of different languages. Like, we didn't take care of having multiple languages, but there's probably, like, there's definitely other languages in there. Yeah, the naming, like, it changed through the project, but we kept it. The idea was also to have four modalities, so it was fitting at the beginning. At the beginning, we wanted audio, video, text, image. We thought that audio and video, video is just a lot of compute for not that much results right now. And from the moment we decided to reproduce Flamingo, we dropped video and audio, which was not immediately.
1:06:42This started as a different project. What's wrong with audio? It's still very, very heavy and not necessary right away because you can take the text tokens, plug it into an audio model after, and it will just spit out the audio. Ultimately, I think it will be good because it's different data as well. Like when people talk, it's not the same as when people write. So I feel there's a lot of interesting data to have with audio. Again, not worth the compute right now, I think. And also you would want this model to be integrating everything with a given size and then be applicable to whatever. You give it to your robot and you're like, what?
1:07:25Yeah. And to close the loop there, Oblix was an asterisk reference that you made into a backonym or whatever. Into an acronym. Yeah, it was really hard to find the acronym. I want to stress this out. I thought about it for days. We were brainstorming it for a few days. We had to cheat at the end a little bit. we took the S of cross-attentions and added it to... Yeah, cross-attentions. The problem is the F of Flamingo, right? For AD6. But we are moving away from the Flamingo architecture. For the next iteration, it wouldn't be really accurate, but whatever. But yeah, the acronym works only for the first version because if I remember well, it's image-aware decoder enhanced a la Flamingo with interleaved cross-attentions.
1:08:20Yes. Yeah, it's a very impressive project. And we were talking about it even before the GPT-4 vision rollout that this is the most impressive sort of open source reproduction. And it's amazing that you're still continuing to work on it. I think a lot of room left in Oblix to keep mining those rocks. and there's a lot of learnings as well in multimodality. I think it's a very important area of research. So thank you. Yeah, thank you.
1:08:55Hello, hello. This is Swix coming in from the editing room in 2024. If you're listening in this far, you're definitely a true fan. Thanks so much. And I hope you enjoyed that conversation as much as I enjoyed recording it. We recorded that conversation on Halloween in 2023 in the hopes that we would be able to release it at NeurIPS or around NeurIPS with IDFX V2. V2 was supposed to be updated with a new base model, which is going to be Mistral, and a bunch of other dataset updates. And a couple of things have happened, but you know what? Hey, let's just get Leo back on to talk about it. Hello from 2024.
1:09:34It's Leo. Just wanted to add a few things since the recording was done a while ago now. We haven't trained the model yet. because we're trading a lot more on data. So we're making progress on OCR, on image-to-code capabilities, and we want to be a lot more thorough in the image text pairs data set that we use. I don't know if you've heard, but the Lay-in data set has had an issue with CSAM images. And so we want to get ahead of this and fix the problem before we start the training. But we will start training soon. In the meantime, we released a website. It's an image to HTML dataset. The idea behind this dataset was to show that we could create a very useful synthetic dataset at scale with open source models.
1:10:25Hugo spearheaded the effort, and so he used Mistral and the DeepSeq coder model to generate the pairs of screenshots and HTML code. and then we fine-tune an early version of IDFX 2.0 on the dataset to have a demo. You can find the model on the hub, you can find the demo on the hub, and you can find the dataset on the hub. So definitely go there and check it out. I think some people have already started training their own model on it. But it's been very well received, and so we think we're going to do a lot more small releases like websites in the future. Basically switching from releasing everything in one package with the model trained to releasing the data sets, architecture and training insights that we get along the way and then releasing the model.
1:11:21So stay tuned. I hope you will enjoy the podcast. And thanks, Sean, for having us.
1:11:39Thank you.
From the publisher
Latent Space is heating up! Our paper club ran into >99 person Discord limits, oops.
We are also introducing 2 new online meetups: LLM Paper Club Asia for Asia timezone (led by Ivan), and AI in Action: hands-on application of AI (led by KBall).
To be notified of all upcoming Latent Space events, subscribe to our new Luma calendar (sign up for individual events, or hit the RSS icon to sync all events to calendar).
In the halcyon open research days of 2022 BC (Before-ChatGPT), DeepMind was the first to create a SOTA multimodal model by taking a pre-existing LLM (Chinchilla 80B - now dead?) and pre-existing vision encoder (CLIP) and training a “glue” adapter layer, inspiring a generation of stunningly cheap and effective multimodal models including LLaVA (one of the Best Papers of NeurIPS 2023), BakLLaVA and FireLLaVA.
However (for reasons we discuss in today’s conversation), DeepMind’s Flamingo model was never open sourced. Based on the excellent paper, LAION stepped up to create OpenFlamingo, but it never scaled beyond 9B. Simultaneously, the M4 (audio + video + image + text multimodality) research team at HuggingFace announced an independent effort to reproduce Flamingo up to the full 80B scale:
The effort started in March, and was released in August 2023.
We happened to visit Paris last year, and visited HuggingFace HQ to learn all about HuggingFace’s research efforts, and cover all the ground knowledge LLM people need to become (what Chip Huyen has termed) “LMM” people. In other words:
What is IDEFICS?
IDEFICS is an Open Access Visual Language Model, available in 9B and 80B model sizes. As an attempt to re-create an open-access version of Flamingo, it seems to track very well on a range of multimodal benchmarks (which we discuss in the pod):
You can see the reasoning abilities of the models to take a combination of interleaved images + text in a way that allows users to either describe images, ask questions about the images, or extend/combine the images into different artworks (e.g. poetry).
📷 From IDEFICS’s model card and blog post
The above demo screenshots are actually fine-tuned instruct versions of IDEFICS — which are again in 9B and 80B versions.
IDEFICS was built by connecting two unimodal models together to provide the multi-modality you see showcased above.
* Llama v1 for language (specifically huggyllama/llama-65b) - the best available open model at the time, to be swapped for Mistral in the next version of IDEFICS
* A CLIP model for vision (specifically laion/CLIP-ViT-H-14-laion2B-s32B-b79K - after a brief exploration of EVA-CLIP, which we discuss on the pod)
OBELICS: a new type of Multimodal Dataset
IDEFICS’ training data used the usual suspect datasets, but to get to par with Flamingo they needed to create a new data set.
Enter OBELICS: “An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents”:
* 115B text tokens
* 141M English documents
* 353M images
These bullets are carefully curated and filtered by going through Common Crawl dumps between FEB 2020 - FEB 2023. We discuss the 2 months of mindnumbing, unglamorous work creating this pipeline:
There’s a lot of mentions of ‘multi-modal' web documents’ which deserves some explanation. We’ll show you instead of tell you:
You can see from this graph that OBELICS ends up outperforming the other image-text pairs dataset (LAION in this case) when stacked head-to-head.
You can view a subset of OBELICS and perform visualizations on them here:
2024 Update: WebSight et al
Most of this interview was recorded on Halloween 2023 at HuggingFace’s headquarters in Paris:
In anticipation of an IDEFICS v2 release. However, several roadblocks emerged, including a notable scandal around CSAM in LAION-5B, which affected all models using that dataset. The M4 team have adopted a strategy of shipping smaller advancements in 2024, and the first ship of the year is WebSight, a dataset of 823,000 HTML/CSS codes representing synthetically generated English websites, each accompanied by a corresponding screenshot (rendered with Playwright). This is intended for tasks like screenshot-to-code workflows like Vercel’s V0 or TLDraw, and will be part of the dataset for IDEFICS-2.
As noted in our Best Papers recap, synthetic data is emerging as one of the top themes of 2024, and the IDEFICS/OBELICS team have wasted no time enabling themselves with it.
Timestamps
* [0:00:00] Intro
* [0:00:00] Hugo, Leo’s path into multimodality
* [0:09:16] From CLIP to Flamingo
* [0:12:54] Benchmarks and Evals
* [0:16:54] OBELICS dataset
* [0:34:47] Together Redpajama v2
* [0:37:12] GPT4 Vision
* [0:38:44] IDEFICS model
* [0:40:57] Query-Key Layernorm for training
* [0:46:40] Choosing smaller vision encoders - EVA-CLIP vs SIG-GLIP
* [0:49:02] IDEFICS v2
* [0:52:39] Multimodal Hallucination
* [0:59:12] Why Open Source Multimodality
* [1:05:29] Naming: M4, OBELICS, IDEFICS
* [1:08:56] 2024 Update from Leo
Show Notes
* Introducing IDEFICS: An Open Reproduction of State-of-the-Art Visual Language Model
* IDEFICS Knowledge sharing memo: technical lessons and mistakes
* OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents
* Papers cited:
* BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
* Barlow Twins: Self-Supervised Learning via Redundancy Reduction
* CLIP paper: Learning Transferable Visual Models From Natural Language Supervision
* Vision Transformers paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
* Flamingo paper: a Visual Language Model for Few-Shot Learning
* April 2022 preprint from DeepMind, blogpost
* VQAV2 paper: Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
* OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge (https://okvqa.allenai.org/)
* MMBench: Is Your Multi-modal Model an All-around Player?
* Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
* Sig-GLIP paper: Sigmoid Loss for Language Image Pre-Training
* Nougat: Neural Optical Understanding for Academic Documents
* MMC4 (Multimodal C4): An Open, Billion-scale Corpus of Images Interleaved With Text
* Dall-E 3 paper: Improving Image Generation with Better Captions
* GPT-4V(ision) system card from OpenAI
* Query-Key Layernorm trick: paper (Scaling Vision Transformers to 22 Billion Parameters), tweet
* EVA-CLIP: Improved Training Techniques for CLIP at Scale
* “We intially explored using a significantly bigger vision encoder (the biggest in open-access at that time) with EVA-CLIP. However, we ran into training instabilities very quickly. To lower the risks associated to the change of vision encoder, we decided to continue with laion/CLIP-ViT-H-14-laion2B-s32B-b79K which we have been using until that point. We will leave that swap for future iterations and will also consider using higher resolution images.”
* Datasets
* Together’s RedPajama-Data-v2: An open dataset with 30 trillion tokens for training large language models
* LAION COCO: 600M synthetic captions from Laion2B-en
* Chip Huyen’s writeup on LMMs
* Joseph Nelson of Roboflow on Latent Space
* HuggingFace timm: library containing SOTA computer vision models, layers, utilities, optimizers, schedulers, data-loaders, augmentations, and training/evaluation scripts. It comes packaged with >700 pretrained models, and is designed to be flexible and easy to use.
* Logan Kilpatrick declaring 2024 the year of Multimodal AI at AI Engineer Summit
Get full access to Latent.Space at www.latent.space/subscribe




