Exploring Open Source Alternatives to GPT-4V: Democratizing AI Language Models

6 Apr 2024 · 13 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Episode Notes: Exploring Open Source Alternatives to GPT-4V: Democratizing AI Language Models

Podcast Overview Title: AI Today Description: AI Today explores advancements in artificial intelligence, discussing breakthroughs, ethical considerations, real-world applications, and the influence of AI on various industries and society.

Episode Summary In this episode, the host delves into open-source alternatives to OpenAI's GPT-4V, exploring how these projects aim to democratize access to advanced AI language models and foster innovation within the community.

Key Topics Covered

  • Introduction to GPT-4V
  • Described as a game changer in AI due to its multimodal capabilities (understanding text and images).
  • Provides practical applications, such as interpreting images and answering context-based queries.
  • Strengths of GPT-4V
  • Effortless handling of tasks that traditional models struggle with.
  • Enhanced ability to provide clear instructions with visual components.
  • Limitations and Critiques
  • Concerns about privacy (e.g., identifying individuals without consent).
  • Issues with bias against certain demographics and the inability to recognize hate symbols.

Open Source Alternatives

  1. LLAVA 1.5
  2. Developed collaboratively by researchers at University of Wisconsin-Madison, Microsoft Research, and Columbia University.
  3. Combines visual coding with an open-source chatbot (Vicuna).
  4. Notable for its scalability, allowing consumer-level hardware applications.
  1. QuenVL
  2. Developed by Alibaba with contributions from Google models (PA, LIX, Palm E).
  3. Licensed to companies with substantial user bases, making it appealing for large-scale applications.
  1. Adept
  2. A startup focused on AI models for web navigation and complex data handling.
  3. Offers an open-source multimodal model similar to GPT-4V.
  4. Designed to engage with the developer community for feedback and improvements.

Challenges and Considerations

  • Performance Limitations
  • LLAVA excels in zero-shot object detection but struggles with complex images.
  • Adept’s model is less geared towards commercial use, focusing instead on community involvement.
  • Ethical Considerations
  • Concerns regarding the potential misuse of AI technologies.
  • Discussion on the importance of responsible development and moderation mechanisms in AI.

Notable Demonstrations

  • LLAVA's ability to detect objects in images and provide contextual insights.
  • Example scenarios, such as analyzing an image of a dog or interpreting unusual scenes from AI-generated content.
  • The model's capability to read and respond to text from images, which raises questions about CAPTCHA security.

Community Reactions

  • Critiques arose over LLAVA's response to an image of a larger woman, highlighting sensitivities related to health and body image.
  • The host defends LLAVA's response as being responsible and contextually appropriate.

Future Outlook

  • The host expresses excitement about the innovation in open-source AI models and the importance of competition in the field.
  • Emphasizes the need for responsible development to prevent monopolization by major players like OpenAI.

Conclusion This episode provides a comprehensive look at the landscape of open-source AI language models, highlighting innovations, ethical considerations, and the importance of democratizing access to advanced AI technologies. The discussion underscores the ongoing evolution within the AI community and the potential for diverse solutions beyond proprietary systems.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The wait is over. Dive into Audible's most anticipated collection, The Best of 2025. featuring top audiobooks, podcasts, and originals across all genres. Our editors have carefully curated this year's must-listens from brilliant hidden gems to the buzziest new releases. Every title in this collection has earned its spot. This is your go-to for the absolute best in 2025 audio entertainment. Whether you love thrillers, romance, or nonfiction, your next favorite listen awaits. Discover why there's more to imagine when you listen at audible.com slash best of the year. OpenAI's GPT-4V has been essentially called a game changer in the realm of AI.

0:45I tend to agree with this. It's pretty impressive as far as all of its multimodal capabilities, which really kind of help it to understand both text and images. And I think this versatility opens up a whole bunch of different applications, but I think it also prevents a bunch of different challenges. So today I wanted to dive into some of the strengths and vulnerabilities of this next, you know, generation of AI and explore some of the open source projects that are also in the same arena. So I can give you, keep you guys up to date on, you know, what else is out there. And I think one other thing, I don't know if I mentioned there in the intro was of course, the fact that GPT-4 now you can upload images.

1:21So I could take an image of like a, you know, a street that has a bunch of parking things on and say, when am I safe to park here right now? It can let me know. you could go take a picture of a menu and say like, hey, what on the menu looks good for someone that's lactose intolerant and vegan and is on the cheaper end of the menu, right? Or whatever, right? And then it could give you an idea. So it can read, it can understand context, very smart, very interesting. So yeah, let's dive into all of the capabilities. So, okay, I've literally just like backspace what I've recorded five times. It's a serious tongue twister.

1:57I'll try to get it right this time. If not, we're just moving forward. But multimodal models, geez, multimodal models, geez, it's so hard to say. In any case, CHI GPT or GPT-4, right? They have a bunch of advantages. So they can, obviously, they're quite effortless in handling tasks that other traditional text or image-based models struggle with. So an example of this would be GPT-4. It can give you like instructions that are a lot clearer with a visual component, right? So if you take a picture of a bike and you're like, hey, like, how do I repair this bike? What pieces are missing? Or what tools do I need to repair this bike?

2:34You send it a picture, it can do this. So of course, these models, they're not just limited to just image recognition. They can also like extrapolate and understand image content to a certain extent. And I think this really kind of opens the door to more nuanced applications, you know, suggesting recipes based on the ingredients found in your fridge might be like a good example, right? You should take a picture of your open fridge. Hey, here's my ingredients. What should I make? So there are some interesting, there's some interesting red flags that some people have raised or concerns that some people have had.

3:05I'll cover them. Some I think are warranted. Some I don't think are as important, but OpenAI initially kind of hesitated to release GPT-4V. That's what it's called, GPT-4V. I think the V is for visual. And they did that, like they kind of hesitated because of concerns that it could be employed to identify individuals and images without their consent, which is, you know, kind of a concern that seems fairly validated. So even after it's released, GPT-4v has a couple issues. Some people are saying that it has the inability to recognize hate symbols or, you know, showing a propensity to exhibit biases against certain genders, demographics, and body types.

3:51It's all the same stuff that, you know, ChatGPT had when it first came out, and then they kind of eventually put up guardrails for it, but different people are complaining, I think, you know, yeah, anyways. So the whole critique is coming directly from inside of OpenAI itself, which is interesting, right? They're the ones critiquing themselves, but of course they launched it because you got to test it, you got to find out how you can break the thing, and I'm sure they've done a ton of internal stuff, but yeah, In any case, so I think despite all of the controversy that it has raised, of course, this is an incredible tool that's incredibly powerful.

4:25So I'm very happy that it has been released. And also, there's a bunch of, you know, both commercial and independent developers that are actively pursuing the development of other open source multimodal models that aim to achieve some similar functionality to this. So some of these are less complex, but I wanted to take a close look at some of them. So there is LLAVA 1.5. So recently, a collaboration between researchers at the University of Wisconsin-Madison, Microsoft Research and Columbia University created this. Essentially, it's an evolution of the original LLAVA, which combines a and so now it just kind of combines a visual coder and Vicuna, which is an open source chatbot to make sense of images and texts.

5:10I think what sets it apart is its scalability. So making it feasible for use on consumer level hardware is kind of why it's special. And I think it will get some adoption because of that. The other one I want to talk about is QuenVL and that's Google's offering. So Alibaba's QuenVL and Google's models like PA, LIX and Palm E are making waves in the multimodal space. So QuenVL in particular, it's licensed to companies with substantial user bases, providing a substantial option for large scale applications. Google, who of course is a major player in AI, is also actively pursuing some of these multimodal models.

5:58So I think this is going to be interesting between those two. The third thing I want to talk about is Adept. So Adept is a startup specializing in AI models for web navigation, and it has introduced a GPT-4V-like kind of multimodal text and image model with a unique twist. It understands, you know, quote unquote, knowledge worker data, like charts and graphs, and this makes it good at kind of handling complex data and manipulation manipulation tasks. So let's talk about some of the challenges and limitations. So LLAVA 1.5 for, you know, with all of its promise is not without some limitations. While it excels at zero shot objects, object detection, and also explaining images, it's not great when faced with complex or crowded images.

6:47So I think it's, you know, it's text recognition capabilities are definitely falling short of GPT-4V, which can be, you know, definitely a blessing in disguise considering GPT-4V's vulnerabilities to malicious text-based manipulations. So, I mean, I don't know, some people say it's a blessing in disguise, but in my opinion, if the technology is worse, it's not a blessing in disguise. You know, you could put your guardrails or whatever else you want on it, but I'd rather the technology to be better and not just say, like, whatever. Yeah, anyways, I think it's a silly argument. But in any case, FUF or FUYU8B, this is what Adept has created.

7:28And that is kind of their, their open source project for this. But it's essentially an open source multi modal model. And it is different in its approach. So it's not designed for commercial use. And it's but you know, yet, maybe they'll do that in the future. And its training data is restricted under specific terms. So instead, Adept is looking to engage with the developer community. They're trying to get feedback and do bug reports. So essentially, this new software out of them has the ability to understand unstructured data, which I think does make it very valuable for knowledge workers. However, it doesn't come with all capabilities prepackaged.

8:05Instead, Adept fine tunes more advanced versions for internal use. And I think this raises questions about like potential misuse and also the need for moderation mechanisms and a bunch of other stuff well yeah i mean that's really a lot of people's concerns about this technology in any case i think you know if we're kind of looking at the road ahead um there are a lot of developers that are pushing boundaries in what open models open source models can do um some people are saying you know the risk and limitations of these models are evident um i think as we kind of move forward on this. It's important that, you know, people obviously develop things responsibly.

8:43But I think that this is a very exciting time where, of course, there's these really powerful tools out of open AI that are closed source and you got to pay for it. But there are a lot of people pursuing this in a more open source manner. And I like to hope that, you know, I'm like not hoping for the demise of ChadCubut and open AI because I do love the technology a lot of the time. But I do hope that there are a lot of other competitors, open source players in this space that do get used and incorporated, I would hate for, you know, the AI space to get sucked up into the vacuum of just the biggest player in the space.

9:14So definitely is exciting to see these other projects. And it's, yeah, it's, it's pretty, it's pretty cool. So one thing I did want to say is LLAVA in testing it, you know, someone uploaded a picture of a dog, it said, you know, detect a dog on the image, return coordinates for this. And it actually like, it sent back the coordinates of where the dog was located, which is kind of an interesting concept to me, right? It's like computer vision, but it's giving you like the coordinates on an image. So you'll be able to select it. That was kind of cool. Another thing that LLAVA's, you know, chatbot and their image detector could do, there was like a picture of a guy.

9:53And this was like an AI generated picture, but it's like a taxi driver and a guy's standing on the back bumper of the taxi and he's ironing his shirt while standing on the back bumper of a taxi. And it says, you know, what is unusual about this image? And LLAVA's chatbot was able to say, the quote, the unusual aspect of this image is that a man is ironing clothes while standing on the back of a moving car. This is not a typical scene is ironing clothes is usually done indoors in a stationary position and with proper safety measures. Anyways, so I thought that was pretty good. Like, obviously, this thing's very intelligent.

10:23That's a kind of response I'd expect to get out of something like Chai Chibi Tea. And what LLAVA has when you ask it a question like that is it has the ability to upvote, downvote, flag, or regenerate or clear. So you are able to, I guess, flag stuff. Another thing that someone tested out on LLAVA, they took a picture of a pile of coins and said, how much money do I have? And it said, you have a total of four coins, which is a total worth of two euros. In addition, they took a screenshot of just like text from someone's like blog post and said, read the text from this image. And it, you know, was able to type out all the text from the image, which I think was very interesting.

11:02You could think of a lot of people might have concerns about like bots essentially doing CAPTCHAs with this technology. Like this would be super easy to write a script to get all CAPTCHAs done in the future, which you've seen CAPTCHAs are definitely getting a lot more complicated now. All sorts of like rotate this animal till it's in the correct position or whatever. So So that'll all get a little bit more, you know, there's going to have to be a lot of work done for CAPTCHAs. One thing that a lot of people were criticizing, if you go to LLAVA and someone did this, I saw a screenshot, but essentially they uploaded a picture of a woman and they said, what advice would you give this person?

11:43And it says, as a responsible AI, I cannot provide personalized advice without knowing the specific context of the person's situation. However, in general, it is important for individuals to remain or to maintain a healthy lifestyle, which includes a balanced diet, regular exercise and proper sleep. What it's what it's, you know, what it's saying here is a picture of a larger woman. And so essentially the researcher was trying to get it to say, like, you know, what advice would you give this person is trying to get to say, like, she needs to lose weight. Right. But I mean, essentially here, I'm not even that mad about its response.

12:14It's just saying, like, I can't give it, you know, I can't give personalized advice without getting known in any context. but like in general balanced diet regular exercise proper sleep these factors contribute to overall well-being and can help manage weight and improve physical health additionally seeking guidance from a health care professional or nutritionist can provide tailored advice and support for achieving and maintaining a healthy weight so a lot of people were triggered by this and were criticizing this because it talks it says the word maintaining a healthy weight at the end of this um yeah so if That's what people are triggered about.

12:45I'm really not mad about this response here. It doesn't seem to go off on any unkind things. It's it's very like I don't know the context. Go talk to a health care professional in general, like healthy lifestyle, balanced diet, regular exercise and proper sleep. Like those are all things that I don't think are very controversial. I know some people will probably criticize me, but whatever. That is what it is. So, yeah, anyways, that was LLAVA. I think like in my opinion, that's probably guardrails making it respond in that way at all. So I personally think that does a great job. It's going to be interesting to see which of these AI models comes out ahead.

13:22There's going to be controversies with all of them. So it'll be a very interesting topic to continue following into the future.

From the publisher

In this episode, we delve into the world of open-source alternatives to OpenAI's GPT-4V, discussing how these projects aim to democratize access to advanced AI language models while fostering innovation and collaboration within the community.

See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

More from AI Today

All 897 episodes
Exploring Open Source Alternatives to GPT-4V: Democratizing AI Language ModelsAI Today · 13 min
Listen in VO