In short
The AI Daily Brief: A Preview of the AI Agent Future
Podcast Overview
- Title: The AI Daily Brief (Formerly The AI Breakdown)
- Description: A daily news analysis show on artificial intelligence covering creativity, industry disruptions, and philosophical questions surrounding AI.
Episode Summary Episode Title A Preview of the AI Agent Future
Episode Highlights
- Adobe's Firefly Image 3 Model Launch
- New generative AI upgrades in Photoshop.
- Firefly Image 3 includes features like:
- Text-to-image generation to overcome the "empty page problem."
- Enhanced capabilities for modifying images using reference images.
- New background generation feature for marketers.
- Apple's Acquisition of DataCollab
- Acquisition of a startup specializing in low-power, high-efficiency deep learning algorithms.
- Indicates Apple's focus on on-device AI capabilities.
- Microsoft's Launch of PHY3
- Introduction of a small AI model that performs comparably to larger models.
- Focus on developing efficient AI systems that can run on devices.
- SoftBank's Investment in AI
- Plans to invest nearly $1 billion to develop a Japanese language-specific AI model.
Main Discussion
The Future of AI Agents Context
- AI agents are gaining traction as they promise to execute tasks autonomously with minimal human input.
- The initial excitement around AI agents (e.g., AutoGPT, Baby AGI) has matured into a focused development phase.
Definition of AI Agents
- AI agents are capable of executing entire strategies, including complex subtasks.
- Examples:
- Instead of researching and booking flights manually, an agent can autonomously handle the entire process.
Current Developments
- Major tech companies (Microsoft, OpenAI, Google) are developing various types of agents:
- Computer-Using Agents: Can operate applications on a user's computer and automate tasks.
- Multi-Step Application Agents: Carry out complex tasks within an application without human oversight.
- Web-Based Task Agents: Complete web-based tasks, such as travel planning.
Incremental Approach to Development
- Companies are focusing on gradual improvements in their existing software rather than launching complex agents immediately.
- Example: Microsoft is developing incremental features in its Dynamics app to suggest actions proactively.
Technical Advances Enabling AI Agents
- Improved use of large language models (LLMs) for generating synthetic data.
- Grounding techniques help verify the validity of AI outputs, enhancing problem-solving capabilities.
Multi-Agent Collaboration
- The concept of using multiple agents to break down complex tasks into subtasks for better efficiency.
- Example: Different agents for software engineering, design, and quality assurance collaborating on a project.
Emerging Concepts
- Payman: A new AI agent tool that envisions agents paying humans for tasks, suggesting a future of collaboration between AI and humans.
Cautions and Critiques
- Some experts caution against over-hyping AI agents, mentioning historical obstacles faced in AI development.
- Concerns about complexity and feasibility remain, but technological advancements suggest a promising future.
Conclusion
- The conversation around AI agents is intensifying, with potential applications extending beyond current AI assistants.
- The trajectory of AI agent development indicates a significant shift in how tasks could be automated and executed in various industries.
Additional Information
- Consensus 2024: The 10th annual event focusing on AI-driven transformation alongside crypto and blockchain.
- Superintelligent Cohort: A platform for learning AI through engaging tutorials, offering a special cohort with access to exclusive resources.
Listening Options
- Subscribe to The AI Breakdown via [YouTube](https://www.youtube.com/@TheAIBreakdown) or the [newsletter](https://theaibreakdown.beehiiv.com/subscribe) for updates on AI developments.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Today on the AI Breakdown, a preview of the AI agent future. Before that on the brief, Adobe releases its latest Firefly Image 3 model, and it looks pretty darn good. The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Go to breakdown.network for more information about our Discord, our YouTube, and our newsletter.
0:24Welcome back to the AI Breakdown Brief, all the AI headline news you need in around five minutes. Since we have launched Superintelligent, a platform for helping people learn AI through fun, fast tutorials, we sometimes have the chance to get previews of new AI software that's coming out a couple days in advance of when the general public sees it. Yesterday, we got to check out Adobe's new Firefly 3 inside of Photoshop. And to give you a sense of how good it is, the first message I got from my tutorial creator who was on the call was that it might even be better than MidJourney. Now, I haven't had enough time to dig into it to make any sort of bold proclamation like that, but the takeaway is certainly that it is a serious contender.
1:00So what are we talking about? Well, Photoshop, the best-known photo editing software in the world for many, many years at this point, many decades at this point, has now gotten a set of generative AI upgrades, including the latest version of Adobe's Firefly Image model, Firefly Image 3. In addition to being embedded in Photoshop, Firefly Image 3 will also be a standalone web app, but it seems likely to me that a lot of people will have their first experiences of it directly inside of Photoshop. Now, for Adobe, they were always going to continue to try to improve their image generation models, but this also really does represent a chance to solve specific problems inside of Photoshop.
1:35In other words, according to Adobe, this isn't just putting in generative AI for the sake of generative AI. Said Zeke Koch, the Adobe VP of generative AI product management, what we noticed when we looked at all these new people who are coming into Photoshop is that they were getting stuck on this empty page problem. So we gave them the ability to generate images, full text to images, because once you have your image in Photoshop, all of a sudden you can use all the tools that we already have in Photoshop to make it better. Now, you might remember that last year, in May, Adobe premiered a feature that they called generative fill.
2:02Basically, users were able to highlight a portion of an existing image and use a text prompt to modify that particular section of an image. An example that they gave in one article is removing a cowboy's lasso from a photo and instead replacing it with spaghetti. This feature has become sort of table stakes for other image generators over the last year, but the new version in Firefly 3 has a bunch of new features that makes it more powerful. One of them, for example, is a reference image where once you've made your selection of the part of the photo that you want to modify, you can use not only text to update the selection, but also a reference image that the model will try to reference alongside your text prompt.
2:35Adobe is also launching something they call generate background, which is exactly what it sounds like. A simpler ability to generate a new background around a particular object of focus. This, as you might imagine, is going to be extremely useful for marketers who need to put their product images in various settings. And this is just the tip of the iceberg. There is a lot more in this product that I, for one, am certainly very excited to dig in on. If you are interested in seeing it in practice, we actually do have tutorials up on Superintelligent right now. You can find them at bsuper.ai. There will also be a code for a discount in the ad that follows the brief.
3:07Next up, however, in the brief, Apple has made yet another quiet AI acquisition. This time it's a Paris-based startup called DataCollab. Writes MacRumors, DataCollab specializes in algorithmic compression and embedded AI systems. The company was established in 2016 and made significant strides in AI technology focusing on low-power, high-efficiency deep learning algorithms that function without relying on cloud-based systems. Obviously, if you've been paying attention to Apple's moves to get their LLMs to run entirely on-device rather than in the cloud, the acquisition just on the face of it makes sense.
3:37So for more information, we once again have to wait until WWDC coming in June. Speaking of small models, Microsoft has launched the PHY3, which is its smallest AI model yet. The Verge writes, the company released PHY2 in December, which perform just as well as bigger models like Llama 2. Microsoft says Fi 3 performs better than the previous version and can provide responses close to how a model 10 times bigger than it can. Now, this is, of course, a reminder that it is not just Apple who is interested in models that can run on device, be it phones or laptops. And even as the race for the state-of-the-art continues on the large model side, all of these companies are also trying to make the most performant small models for these different types of use cases.
4:15Lastly, reports that SoftBank is set to invest nearly a billion dollars in their own AI push. One of the trends that we're seeing right now is companies racing to build highly performant LLMs that are built around non-English languages. CNBC writes that SoftBank is looking to develop a world-class Japanese language-specific artificial intelligence model and plans to spend about a billion dollars on the compute to do so. This is the type of trend that I expect to continue in basically every major language market. For now, though, that is going to do it for today's AI Breakdown Brief. Next up, the main AI breakdown.
4:46Attention, AI Breakdown listeners. Consensus 2024 marks the 10th gathering for all things crypto, blockchain, and Web3. However, importantly, this year's agenda will also dive deep into AI-driven transformation. And the speaker lineup includes the leading minds and innovators at the forefront of this digital renaissance. Don't miss the Consensus AI Summit to cut through the hype to find where true transformation and opportunity lie. Listeners to this show can get 15 % off registration with the code AIBreakdown. Visit consensus2024.coindesk.com to learn more. Some of the folks who will be at Consensus this year include Guillaume Verdon, aka Beth Jezos, founder and CEO of Xtropic, as well as spiritual leader of the accelerationist movement, Neil Stephenson, co-founder of Lamina One, and Brendan Eich, the CEO of Brave Software.
5:30Again, go to consensus2024.coindesk.com to learn more and get 15 % off registration with the code AIBreakdown. Before we get back to the AI Breakdown, I want to share something fun we have coming up on Superintelligent next month. Super Intelligent is, of course, our new platform for teaching people how to use AI in a way that is much more fun, fast, and practical. The platform has hundreds of short tutorial videos, each of which is paired with a set of step-by-step instructions that get you using AI tools in minutes, not hours, and certainly not days. For those of you who haven't signed up yet but want to check it out, in May, I am running a special NLW cohort.
6:04What this means is that people who sign up with the code NLWMay will get$5 off their first month, but they'll also have access to a private channel in our Discord with me. I'll be handpicking tutorials each week that I think are the most useful to start with, and I'll also be available for questions, advice, and feedback from this group. Spots for this cohort are limited, so if you want to be a part of it, again, sign up at besuper.ai with code NLWMay. That's besuper.ai with code NLWMay. Welcome back to the AI Breakdown. Since very shortly after the launch and explosion of ChatGPT, AI builders have been extremely excited about AI agents.
6:45You might remember back around this time last year when things like AutoGPT and Baby AGI were capturing everyone's imagination. Simply put, the thing that has people excited about an AI agent future is that instead of using AI as a co-pilot to crib the term that Microsoft has laid claim to, or an assistant, AI agents are capable of executing a strategy from end to end, including subtasks. To take a very banal example, instead of using AI-supported research to find the best deal on flights, but then booking them yourself, you would simply tell an agent to buy the best flight based on whatever criteria you had, and it would be able to actually do that.
7:20After the initial burst of hype and excitement around last April-May, the conversation ebbed again. There just wasn't all that much that agents could do yet, and frankly, it was just a quiet behind-the-scenes kind of building time. Certainly there was no shortage of people working on these issues. Towards the end of the year, it became clear that it wasn't just going to be startups who were thinking about agents, but all of the big AI labs. Indeed, Sam Altman hinted at this at OpenAI's Dev Day last year, calling the custom GPTs their very, very first tiny baby steps towards that agent future. Well, the information has this week done a report that serves as kind of a check-in on the state of the agent conversation.
7:56The piece is called To Unlock AI Spending, Microsoft OpenAI and Google Prep Agents. The gist of the piece is that as the big AI labs try to convince enterprises to spend more on AI, they're making a bet that agentic software is likely to pry open those wallets even better than what's available now. The information writes, Microsoft is making software to automate multiple actions such as creating, sending, and tracking a client invoice based on their order history or rewriting an application's code in a different language and verifying that it works as intended. This is just one example that the piece gives.
8:26They continue, The features belong to a class of AI software known as agents, bots that can work towards a goal with minimal guidance from people. OpenAI, Google, and Facebook owner meta-platforms are each developing their own versions of agents. It's part of the industry's broader effort to turn the excitement ChatGPT sparked 18 months ago into recurring revenue for a slew of companies that sell such technology. While AI chatbots have wowed the business world with their ability to generate realistic answers or suggest lines of code to a programmer, customers say software that automates harder tasks will be necessary to unlock more spending.
8:54So what are some of the other types of agents that they say these companies are working on? Well, they write, OpenAI is quietly designing computer-using agents that could take over a person's computer and operate different applications at the same time, such as transferring data from a document to a spreadsheet. OpenAI and Meta are also working on a second class of agents that can handle complex web-based tasks, such as creating an itinerary and booking travel accommodations based on it. The way that the information sums up different agent types includes computer-using agents, which we just discussed, something that could take over a user's computer, including moving the cursor and keyboard and using different applications.
9:25The sample task they give is autonomously conduct research across a user's file and online sources and compile a new presentation. A second type of agent they mention is multi-step application agents. The idea here is an agent that can carry out multiple step tasks within an application without human oversight between steps. So an example is something like Microsoft is apparently working on that could draft invoices for customers using data inside a company's sales software than record and summarize customer payments. A third category of agents they call web-based task agents. As you might guess, these complete web-based tasks that require communicating with different applications, and a sample task could be researching and planning a user's vacation.
9:58Now, in terms of how these roll out, a lot of people seem to think that the right approach is very incremental. For example, again from this piece, instead of launching the most sophisticated form of agents, some companies like Microsoft are looking to launch ones that incrementally improve the automation features in current versions of its software. Earlier this year, Microsoft formed a new team under Scott Guthrie, executive vice president of cloud and AI, to develop agent capabilities for the company's co-pilot products. An upcoming agent feature Microsoft is building within its Dynamics app for salespeople, for instance, aims to proactively suggest multi-step actions the app can take, actions that users previously would have needed to instruct co-pilot to do.
10:31This incremental approach, I think, makes sense as a big part of what was so disappointing to some last year was that when push came to shove, the agent softwares like AutoGPT really couldn't do all that much on their own. They needed a ton of human interaction and engagement. Given that, being extremely specific about these incremental tasks that agents can do that automate one specific type of workflow would seem to me to be a good wedge in while not overpromising. There are also technical advances that are potentially unlocking new agent use cases. One, according to Ion Stoica, a co-founder of AnyScale and Databricks, is that, quote, developers have collectively gotten better at using LLM to generate synthetic data.
11:09That's especially helpful in code generation where developers can instruct models to create and then solve problems within a set of parameters. A second advancement is called grounding, the process of setting up AI models that can automatically verify whether another model's outputs are valid, such as by testing whether the code a model generated solved the problem at hand correctly. Stoica continued, In the coming year, we're going to see a significant jump in the model's ability for problem solving and reasoning. This will hinge on grounding. If I can automatically verify that an output is valid, then I can use LLMs themselves to improve the outputs, which is huge.
11:38Now from there, I wanted to share a post on X recently from Andrew Ng, the co-founder of Coursera, and the former head of AI at Baidu and Google Brain. He writes about something similar, saying, Multi-agent collaboration has emerged as a key AI-agentic design pattern. Given a complex task like writing software, a multi-agent approach would break down the task into subtasks to be executed by different roles, such as a software engineer, product manager, designer, quality assurance engineer, and so on, and have different agents accomplish different subtasks. Different agents might be built by prompting one LLM or if you prefer different LLMs to carry out different tasks.
12:11For example, to build a software engineer agent, we might prompt the LLM, you are an expert in writing clear efficient code, write code to perform the task, dot dot dot. It might seem counterintuitive that, although we are making multiple calls to the same LLM, we apply the programming abstraction of using multiple agents. I'd like to offer a few reasons. First, it works. Many teams are getting good results with this method, and there's nothing like results. Further ablation studies showed that multiple agents give superior performance to a single agent. Next, even though some LLMs today can accept very long input contexts, for example, Gemini 1.5 Pro accepts 1 million tokens, their ability to truly understand long, complex inputs is mixed.
12:45An agentic workflow in which the LLM is prompted to focus on one thing at a time can give better performance. By telling it when it should play software engineer, we can also specify what is important in that subtask. For example, the prompt above emphasized clear, efficient code, as opposed to, say, scalable and highly secure code. By decomposing the overall task into subtasks, we can optimize the subtasks better. Perhaps most important, the multi-agent design pattern gives us, as developers, a framework for breaking down complex tasks into subtasks. When writing code to run on a single CPU, we often break our program up into different processes or threats.
13:15This is a useful abstraction that lets us decompose a task, like implementing a web browser, into subtasks that are easier to code. I find thinking through multi-agent roles to be a useful abstraction. In many companies, managers routinely decide what roles to hire and then how to split complex projects, like writing a large piece of software, preparing a research report, and a smaller task to assign to employees with different specialties. Using multiple agents is analogous. Each agent implements its own workflow, has its own memory, and may ask other agents for help. So given that that is theoretical, I thought it was also interesting that right around the same time, tech evangelist Robert Scoble retweeted a post from Taskade that shows this multi-agent approach being put into practice.
13:51The company writes, Introducing Taskade's multi-AI agents, now entering beta. Imagine one AI agent researching while another converts insights into tasks. They can write articles, perform research, summarize findings, and edit content all at once. Now, I haven't had a chance to see how Taskade works yet, and you can't really throw a rock at Silicon Valley right now without hitting an AI agent startup. But still, it's interesting that we're seeing some of this multi-agent approach go into practice. Other explorations in the AI space are also capturing attention. Zero X Thailand recently got X buzzing with a project he announced called Payman.
14:21Payman, he writes, is an AI agent tool that gives agents the ability to pay people for tasks they cannot do themselves. He continues, While many people imagine a future where humans pay AI agents for services they want completed, I believe that as AI agents become more advanced, they will be paying humans for tasks they can't do. There will always be important roles for humans, and as we move towards an agent-driven world, Payman's goal is to support a symbiotic relationship between AI agents and humans. So the idea here is exactly as he describes, giving agents who are coordinating a set of tasks the ability to coordinate humans as part of that larger picture.
14:51Now, one caveat to all of this comes from Pedro Domingos, who writes, The next big thing in AI is agents, except it's a decades-old idea with whole conferences dedicated to it. It always hits a wall of complexity and goes nowhere, and no one has explained what will be different this time. Now, I think there is a lot different this time in terms of technological capacity, in terms of the sheer amount of energy being poured into this, in terms of the number of different experiments, in terms of the specificity of experiments, but the caution is still well received. Overall, it does seem very clear to me that the AI agent conversation is increasing.
15:21And it doesn't seem to me that it's just because, as the information posits, these big labs are looking for enterprises to spend more. In short, I think it's happening because AI agents are what's next. They're a seemingly obvious extension of some of the workflow automation and workflow reimagining that we've been doing with the assistant-type tools we have now. It is not a foregone conclusion to me that every process will be agentized. And I certainly think that initially, the most success is going to be found with very discrete processes like we discussed earlier. But I am not cynical about what people are trying to build.
15:51I think we have a long way to go, but I think these things are going to be incredibly useful. And so I'm excited to see what people build next. For now, though, that is going to do it for today's AI Breakdown. Until next time, peace.
16:12you
From the publisher
Explore the rapidly evolving world of AI agents in this episode of AI Breakdown. Discover how Microsoft is spearheading projects to automate complex tasks such as client invoicing, while OpenAI advances in enabling desktop automation with minimal human input. This discussion highlights the potential for AI agents to dramatically increase enterprise investment and revolutionize the software industry. Insights into practical and theoretical developments reveal a future where intelligent, autonomous systems handle intricate tasks, paving the way for significant advancements in technology and business efficiencies.
**
Join NLW's May Cohort on Superintelligent.
Use code nlwmay for 25% off your first month and to join the special learning group. https://besuper.ai/
**
Consensus 2024 is happening May 29-31 in Austin, Texas. This year marks the tenth annual Consensus, making it the largest and longest-running event dedicated to all sides of crypto, blockchain and Web3. Use code AIBREAKDOWN to get 15% off your pass at https://go.coindesk.com/43SWugo
**
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI.
Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe
Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown
Join the community: bit.ly/aibreakdown
Learn more: http://breakdown.network/
