Google's AVIS: Revolutionizing AI Image Search

20 Mar 2024 · 6 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI Today Podcast Episode Notes

Episode Title

Google's AVIS: Revolutionizing AI Image Search

Episode Overview In this episode, the hosts discuss Google's latest innovation in AI image search, known as AVIS. This technology enhances image recognition accuracy and fundamentally changes how users search for images online.

Key Topics Covered

  • Introduction to AVIS: The significance of Google's AVIS in image search and how it addresses existing limitations in AI models.
  • Visual Information Seeking: The challenge of obtaining detailed information from images that are not immediately apparent to users.
  • Comparison with Existing Technologies: Understanding how AVIS differs from traditional image search tools like Google Images and Amazon Lens.

Main Components of AVIS

  1. Three Pillars of AVIS Architecture:
  2. Planner: Utilizes large language models (LLMs) for action planning and reasoning.
  3. Working Memory: Archives data from previous API calls for enhanced context and continuity in responses.
  4. Reasoner: Analyzes API results to distill key information and improve response quality.
  1. Dynamic Approach:
  2. AVIS adapts its actions based on real-time feedback, allowing for more nuanced queries and responses.
  3. This iterative mechanism continues until the reasoner is satisfied with the amount of data available to produce a final answer.
  1. Integrated Tools:
  2. Computer Vision: Extracts visual data from images.
  3. Web Search Utility: Accesses open-world knowledge and facts.
  4. Image Search Feature: Mines metadata from similar images to provide context and details.

User Study Insights

  • A user study was conducted to capture human-driven choices while using visual reasoning utilities.
  • Results led to the development of a transition graph that guides the operations of AVIS.

Performance Metrics

  • Accuracy Benchmarks:
  • Achieved 50.7% accuracy on the InfoSeq dataset, outperforming fine-tuned visual language models like OFA and Pali.
  • Logged 60% accuracy on the OKVQA dataset, coming close to the performance of more advanced models.

Future Directions

  • AVIS is currently in the research phase and not yet released as a product.
  • Future applications may explore the efficacy of lighter language models for these functionalities, especially considering the current model, Palm, has 540 billion parameters.
  • The research aims to apply this framework to a wider range of reasoning challenges in AI.

Conclusion

  • The hosts express excitement about the potential impact of AVIS on image search and visual reasoning capabilities.
  • They emphasize the need for ongoing evaluation of AVIS as it develops and its implications for future AI technologies.

Additional Resources

  • Invest in AI Box: [AI Box Investment](https://republic.com/ai-box)
  • AI Box Waitlist: [Join Waitlist](https://aibox.ai/)
  • AI Facebook Community: [Join Community](https://www.facebook.com/groups/739308654562189)
  • AI in Music: [Learn More](https://musicalai.pro/)
  • AI Models: [Explore AI Models](https://aimodelspro.com/)

Privacy Information

  • For privacy policy details, visit: [Privacy Policy](https://art19.com/privacy)
  • California Privacy Notice: [California Privacy Notice](https://art19.com/privacy#do-not-sell-my-info)

---

This markdown file serves as a comprehensive summary of the podcast episode, capturing the essence of the discussions on Google's AVIS and its implications for the field of AI image search.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Essentially, in the expanding world of image-based queries, there are a lot of times when you're going to ask yourself questions about an image's details that are not immediately. apparent. So an example of this would be, you know, taking a picture of an airline and trying to figure out what year this airline was started or taking a picture of a car and deciding, you know, what year this car was manufactured just by looking at its picture. This is not something that you currently can do with AI models and Google thinks that is a problem. So Google, who is the king of search, is essentially creating something to allow you to do this.

0:31So today on the podcast, we're going to be diving into how Google Avis is approaching this field and why we think this is significant. While the progress in LLMs has unlocked multi-model functions such as visual question answering and image captioning, there remains some hurdles where these visual language models, or they're called VLMs, essentially tackle intricate visual reasoning that leans on external information. So this particular challenge is termed visual information seeking. So recognizing the gap and the struggle that a lot of current technology faces, Google researchers have crafted a new method.

1:09It's called Avis, which is a meld essentially of Google's palm, computer vision, and the combination of web and image search utilities. Now, you'll kind of get a general idea of where their background is with this because Google has had essentially a Google image search function where you can take a picture with your camera of an item and it's able to kind reverse google image search it and find you know what that is and also you see this in other technologies for example amazon has an amazon lens feature where you can take a picture of a product and it will find that product on amazon so this kind of image identifier technology is there but now we're we're essentially merging this with more ai technology to answer more in-depth and more intricate questions so distinct from older systems that kind of brought together large language models with tools in a predefined sequence, AVIS employs them in a more adaptive manner.

2:03So emphasizing both planning and reasoning. This dynamic approach essentially ensures actions evolve based on instant feedback. So breaking down its architecture, AVIS is constructed around three main pillars. The first one is a planner that leverages, you know, the LLM to kind of pinpoint the the actions which include an API call and its query. The second component of this is a working memory that essentially archives data from previous API implementations, which is really interesting. And the third one is a reasoner, which taps into the LLM to distill the essential details from API results, right?

2:44So it's actually able to get the results and then run it through a reasoner, which is able to make these, which is able to kind of make them more digestible and better. So this iterative mechanism involves the planner zeroing in on the next query and tool, informed by the latest insights from the reasoner, continuing until the reasoner believes enough data is available to produce the final response. So here is, I guess, moreover how Avis kind of encompasses three integrated tools. So they have tools harnessing computer vision to cull visual data from photos. They have a web search utility to open source, or I guess open world knowledge and facts.

3:25And then they have an image search feature that mines metadata from visually akin images for pertinent details, right? So this is really, really interesting. They're essentially using the Google lens thing that I mentioned earlier, where it's able to take a picture and reverse search it. They're able to reverse search it, find the images that are similar, Then they go to the metadata of those images and they find a whole bunch of them so they can essentially pinpoint what exactly the image is of. It's really interesting and really smart because Google has access and has essentially indexed, you know, millions and billions of images around the web.

3:58And so they have this massive data set that they can kind of harness in this way. So to maximize these unique functions, the Google team essentially went on a user study where they were capturing human-driven choices when utilizing visual reasoning utilities. So this experiment highlighted reoccurring action sequences leading to the creation of a transition graph that instructed Avis's operations. So furthermore, the results are super, super promising. In benchmarks against the InfoSeq dataset, Avis posted an impressive 50.7 % accuracy, which is definitely overshadowing fine-tuned visual language models like OFA and Pali.

4:43And when tested against the OKVQA dataset, it logged a 60 % accuracy with minimal examples surpassing much of the past work and nearing fine-tuned models as per Google. so looking ahead I think the research you know this is not a product they've released yet obviously I mean we're looking at 50 % accuracy 60 % accuracy some people are like oh my gosh that's far from you know accurate that's the flip of a coin or something I mean it's not really a flip of a coin because you're you're getting like 50 % of the time you're getting some very detailed in-depth knowledge about a certain item or picture that you couldn't get any other way so um it is quite impressive but of course not perfect so this is research at the moment and essentially the research it kind of is aimed to apply this framework to a broader array of reasoning challenges.

5:32And they're also curious to find out if lighter language models might execute these functionalities. And this kind of curiosity is essentially pertinent, given that the current used Palm model is a computational giant. This thing has a staggering 540 billion parameters, so it is absolutely massive. And if they're able to essentially use some of these tools, right, where they have different elements of the tool. They got the planner, the working memory, the reasoner, right? They had these different components and they run it through a sequence to get what the output is. And it's much, much more efficient than using something like a 500 billion parameter model.

6:14So it's gonna be very interesting to see how this research and this technology plays out into new tech and new AI models that are coming out. And if this is something that people begin to adopt, Definitely something we'll continue to follow in the future.

From the publisher

In this episode, we explore Google's latest breakthrough in AI image search with AVIS, discussing how it enhances image recognition accuracy and revolutionizes the way users search for images online.

See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

More from AI Today

All 897 episodes
Google's AVIS: Revolutionizing AI Image SearchAI Today · 6 min
Listen in VO