The Most Interesting AI Research This Week

18 Jun 2023 · 12 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The AI Daily Brief: Episode Summary

Episode Title

The Most Interesting AI Research This Week

Podcast Overview

  • Podcast Title: The AI Daily Brief (Formerly The AI Breakdown)
  • Description: A daily news analysis show that covers various aspects of artificial intelligence, from creativity and industry disruptions to philosophical and ethical questions surrounding AI.
  • Host: NLW (Nathaniel L. Weller)

Episode Description This episode presents a research roundup highlighting significant developments in AI, focusing on:

  • AssistGPT
  • LLaMA multimodal adapter
  • Meta Voicebox
  • Text-to-Video technologies
  • LLM agents teaching weaker AIs
  • AI art QR code generator

Key Themes Discussed

  1. Auto-GPTs and AI Agents
  2. Introduction of AutoGPT by Sig Gravitas, which became a major project on GitHub.
  3. Other related projects such as BabyAGI aimed at creating AI agents that can perform tasks and even develop other AIs.
  4. AssistGPT:
  5. Not a direct competitor of AutoGPT but integrates multimodal capabilities.
  6. Follows the Plan, Execute, Inspect, Learn (PEIL) framework:
  7. Planner: Uses natural language to strategize.
  8. Executor: Carries out planned tasks with potential for visual processing.
  9. Inspector: Manages memory and provides necessary visual data.
  10. Learner: Improves performance through experience.
  1. Text-to-Video Technology
  2. New framework for transforming text prompts into videos without flickering issues.
  3. Keyframe Translation: Generates keyframes from text; ensures coherence in visuals.
  4. Full Video Translation: Fills gaps between keyframes for continuity.
  5. Implications for creative industries and content generation are significant.
  1. Meta Voicebox
  2. AI capable of high-quality audio generation in six languages without specific task training.
  3. Utilizes flow matching for improved speed and intelligibility.
  4. Versatile with applications in speech synthesis, style conversion, and noise removal.
  5. Meta has withheld public release due to potential misuse but has shared samples in research papers.
  1. Robotics and Object Catching
  2. Research on robotic capabilities for high-speed object catching through a combination of Model Predictive Control (MPC) and Reinforcement Learning (RL).
  3. Potential commercial applications, enhancing robots' ability to perform tasks traditionally done by humans.
  1. Teaching AI Models
  2. Study on whether advanced LLMs (Large Language Models) can teach weaker AIs.
  3. Establishes a teacher-student framework to explore effective explanation methods.
  4. Findings indicate that teacher AIs improve student performance, with implications for AI training and alignment.
  1. Multimodal Learning with Llama Adapter
  2. Llama Adapter V2 integrates visual and textual data, enhancing instruction-following capabilities.
  3. Promotes interactive applications such as chatbots and educational tools.
  1. AI Art and QR Codes
  2. Rise of AI-generated QR code art on social media platforms.
  3. Hugging Face has developed a QR code AI art generator, merging visual creativity with digital information.

Conclusion The episode encapsulates a week of innovative AI research, showcasing advancements in multimodality, creative AI applications, and the potential of AI agents to revolutionize tasks across various sectors. The discussion emphasizes the dynamic nature of AI development and the importance of ethical considerations in its deployment.

Additional Resources

  • Subscribe to the [AI Breakdown newsletter](https://theaibreakdown.beehiiv.com/subscribe)
  • Watch on [YouTube](https://www.youtube.com/@TheAIBreakdown)
  • Join the community: [AI Breakdown Community](bit.ly/aibreakdown)
  • Visit the [Breakdown Network](http://breakdown.network/) for more insights.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Today on the AI Breakdown, we're looking at the most interesting research from the previous week. Big themes like multimodality and text-to-video are all over this thing. The AI Breakdown is a daily podcast and video about the most important news and stories in AI. Like, subscribe, and share, and go to breakdown.network for more information. Today, we are looking at some of the most interesting AI research from the last week, and we're kicking off with a theme that has really defined the last few months, and that is auto-GPTs, and more broadly speaking, the interest that people have in AI agents that can actually go out and do tasks.

0:36Now, if you've listened to this show or watched this channel for any amount of time, you have inevitably heard of AutoGBT. In April, Sig Gravitas introduced AutoGBT, and it quickly became the most interacted with project on GitHub. Other projects like it include BabyAGI, and basically, again, these tools were all designed to try to be an approach to an AI agent that could actually not only figure out how to solve problems, but to potentially spin up and create the other AI agents that could go out and do the tasks that were necessary to accomplish a particular goal. AssistGPT isn't exactly a one-to-one auto-GPT competitor or anything like that.

1:10Instead, it's an approach that tries to bring in another theme of this year, which is multimodality or large language models that can also interact with visual-based tasks, such as images, videos, and more, to try to create a more sophisticated AI assistant. Assistant GPT executes tasks using an approach called Plan, Execute, Inspect, and Learn, P-E-I-L, which is what has reminded some people of Auto-GPT. The planner uses natural language to plan the next steps based on the current progress of reasoning. It decides which tool or function should be used next to process the input data or move closer to the final answer.

1:42The executor carries out the tasks planned by the planner, which could involve processing visual data such as images or video, or it could involve other types of tasks such as searching for information or performing calculations. The inspector is a memory manager that helps the planner by providing the appropriate visual information to the specific tool that needs it, which again could involve selecting the right images or videos or involve providing other types of data that the tool needs to perform its task. And the learner is designed to improve the system's performance over time. It allows the model to learn from its experiences and to discover the optimal solution to a task.

2:11This learning process could involve adjusting the way the planner, executor, and inspector work based on the results of the previous tasks. Now, the paper doesn't specifically mention the ability of AssistGPT to interact with or create other AI agents to execute tasks. However, the concept of AssistGPT involves integrating LLMs with various tools, which could potentially include other AI models or agents. Next up, we have re-render a video. This is one that people like Matt Wolf got really excited about. So as Matt puts it, it takes an input video and re-renders it with your prompt without the flicker or weirdness we get from other current models.

2:45For those of you who are listening or not watching, you can see how it takes a source video and then can re-render it based on a variety of different styles. So a colorful impasto painting, a Ghibli cartoon, a starry night of Van Gogh. Basically, all of these text prompts change the nature of the video to look like the text prompt. This research presents a novel framework for translating text into video. The framework consists of two parts, keyframe translation and full video translation. Keyframe translation uses a modified diffusion model to generate keyframes from text with constraints applied to ensure coherence in shapes, textures, and colors across frames.

3:21And then full video translation fills in the gaps between keyframes using a method that matches and patches blends from the keyframes, which ensures style and texture consistency over time. One of the big things we've seen over the last couple months is that all of that excitement and energy that came to the text to image space last year coming into this year, right? The difference between mid-journey two and mid-journey five that we have now is starting to make it to video as well. Runway Gen 1 had this sort of video reskinning capacity and then Runway Gen 2 has an entire video generation suite where you can actually go text to video in four second snippets.

3:56The creative, artistic, and business implications of this sort of text to video approach are enormous and people are kind of salivating to get their hands on anything that can do this type of video modification or generation. Next up we have Meta's VoiceBox. As Min Choi sums it up, it's an AI that can create high-quality audio in six languages, with noise removal, content editing, style conversion, and more, all without specific training. Min writes, unlike traditional speech AI that needs specific training for each task, VoiceBox learns from raw audio and transcriptions. It's based on flow matching and outperforms current models in speed, intelligibility, and audio similarity.

4:33Unlike old models that require carefully prepared data, Voicebox can learn from varied speech data without needing careful labeling. It's trained on diverse data at a much larger scale, making it versatile and adaptable. Voicebox was trained on 50 ,000 hours of recorded speech and can perform tasks including text-to-speech synthesis, cross-lingual style transfer, speech denoising and editing, and more. Now there are just a ton of potential use cases for this. With just a two-second long audio input sample, Voicebox can match the sample's audio style and use it for text-to-speech generation. That could potentially be used to provide speech capabilities for people who are unable to speak, to allow individuals to customize the voices used by non-player characters or virtual assistants, to have themselves saying things in a different language in the same style as they normally would in whatever their home language is.

5:17And again, it could also be totally transformational in how fast and easy it is to clean up audio, taking out background noises or imperfections and things like that. This is Voicebox, a new AI foundation model for speech that does some pretty awesome things. If you give it text, it can read it in a bunch of different styles. Penelope Porcupine and Sammy Sloth danced gracefully in the treetops. Penelope Porcupine and Sammy Sloth danced gracefully in the treetops. And you can use it to fix background noise too. Kind of like an eraser, but for audio. Sammy and Penelope's heartwarming friendship inspires joy.

6:03We think this is probably the most versatile speech generative model out there. This is still a research project, but I think that we're going to be able to build a lot of interesting things with tools like this. Meta puts out so many different AI models at this point that it's hard to keep track of them all, but this one really does strike me as something pretty significant. They claim it as the first ever generative AI speech model that can do tasks it wasn't specifically trained on. That is exactly the sort of transformation that led to all of the tools that we have today when it comes to things like text-to-image.

6:34Fascinatingly, however, Meta has decided not to release the VoiceBox model or code publicly at this time due to the potential risks of misuse. As part of their efforts to mitigate that potential misuse, they've developed a classifier that can distinguish between authentic speech and audio generated with VoiceBox, but they still decided to share audio samples in a research paper instead of the actual code. For those of you who aren't worried about AI alignment until robots get bodies that are human-like, I have some bad news for you. A new research paper titled Agile Catching with Whole Body MPC and Black Box Policy Learning is all about a new approach to enabling robots to catch objects thrown at high speeds.

7:10The researchers here explored two different solution strategies, including model predictive control, MPC, using accelerated constrained trajectory optimization, and reinforcement learning, RL, using zero-order optimization. Now, TLDR, the combination of these methodologies, has led to some really big advancements. This sort of high-speed object catching is an extremely complex task. It requires split-second decision-making and precise mechanical control. There are, of course, huge commercial implications, giving robots the ability to do things where fast and accurate object handling is crucial, or, you know, robots might just become the goalies of the future.

7:42And speaking of the robots now being able to do a thing that only humans used to do, Another interesting piece of research this week is a paper titled, Can Language Models Teach Weaker Agents? Teacher Explanations Improve Students Via Theory of Mind. In simple terms, the researchers were trying to understand if advanced AI models, specifically LLMs, can act as teachers to less advanced AI models, which are referred to here as weaker agents. They're interested in whether these LLMs can improve the performance of weaker agents by providing them with explanations for their predictions. The researchers set up a student-teacher framework where the LLM, the teacher, provides explanations to the weaker agent, the student.

8:16However, they also set a limit on how much the teacher can communicate with the student to mimic real-world constraints where resources like time and computational power might be limited. The researchers explore four main questions. Can the teacher's intervention improve the student's predictions? When is it worth explaining a data point to the student? How can the teacher personalize explanations to better teach the student? Can the teacher's explanation improve the student's performance on future data that hasn't been explained? TLDR, the researchers, found that teacher LLMs can indeed improve the performance of the student agents.

8:43They also developed a theory of mind approach where the teacher builds two mental models of the student. The first model helps the teacher decide when to intervene, and the second model helps the teacher personalize explanations. Now, this obviously has huge implications for how we train future AI systems. It could be used not only for efficient AI training, but also help personalize AIs for specific use cases. It could be applied in scenarios where multiple AI systems need to work together, where the teacher AI could help improve the performance of the student AI, leading to a more effective collaboration.

9:12And there are some interesting implications for AI alignment. The teacher-student framework could be used to train AI systems to better understand and mimic human values, where the teacher could be a model that has been trained to understand human values and then pass on that understanding to the student. It could personalize alignment for different quote-unquote students. But it also showed that this sort of AI alignment is important, as the study found that misaligned teachers can lower student performance by intentionally misleading them. One more on the theme of multimodality, another big theme of the year.

9:40Remember, OpenAI thought that they were going to bring a multimodal model to ChatGPT in 2024. In other words, a model that can interact with text, image, audio, and video inputs, not just text inputs, but have been constrained by GPU access. Well, this new research, titled Llama Adapter V2 Perimeter Efficient Visual Instruction Model, is an approach to multimodality that works with the Meta Llama model. They made more parts of the Llama model learnable, which means the model can adapt better to the task of following instructions. They introduced a new strategy for incorporating visual information into the model.

10:11Instead of feeding visual information into all the layers of the model, they only feed it into the early layers. This helps the model incorporate visual knowledge better. They train the model on two types of data, image text pairs and instruction following data, which helps the model get better at both understanding images and following instructions. And during the use of the model, they incorporate additional expert models like systems that can generate captions for images or recognized text in images to enhance image understanding capabilities. They found that the new model, Llama Adapter V2, is better at following instructions that involve both text and images and even performs well in chat interactions.

10:42So the research opens up new possibilities for AI applications, such as more interactive chatbots or educational tools. And TLDR, this is just one of the big themes going on in the space right now. Just to show a super quick example, let's upload an image of a Bernice Mountain Dog, as you can see, and type what breed of dog is in the picture and then run it. Around 10 seconds later, the dog in the picture is a Bernese Mountain Dog. Now, these guys are obviously far from the only team working on multimodality, but cool to see that research live and demos live as well. Last up, one of the biggest viral trends over the last week has been QR code art.

11:19AI-generated visual QR codes started popping up all over Reddit and then Twitter over the last week, but then the good people at Hugging Face went and put together a QR code AI art generator. Radamus Ajna says you only need the QR code content and a text to image prompt idea, or you can upload your image. This is one that is really better seen than described. So if you are listening to this, I suggest you check out the YouTube video as well. All right, guys, that's going to do it for today's AI breakdown. Hope you enjoyed this. Thanks to so many great research teams for sharing their knowledge.

11:50Exciting to see what's coming down the line. If you're enjoying the AI breakdown, please like, subscribe and share it. Check out the podcast version and the newsletter version. and until next time, peace.

From the publisher

A Research Roundup including: -AssistGPT -LLaMA multimodal adapter -Meta Voicebox -Text-to-Video -LLM agent teaching weaker AIs -AI art QR code generator   The AI Breakdown helps you understand the most important news and discussions in AI. 
Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe
Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown
Join the community: bit.ly/aibreakdown
Learn more: http://breakdown.network/

More from The AI Daily Brief: Artificial Intelligence News and Analysis

All 1,099 episodes
The Most Interesting AI Research This WeekThe AI Daily Brief: Artificial Intelligence News and Analysis · 12 min
Listen in VO