In short
Podcast Summary: The AI Daily Brief - OpenAI Building AI Agents as Google Launches Gemini Advanced
Episode Overview In this episode of The AI Daily Brief, host NLW discusses significant updates in the AI landscape, focusing on:
- The transformation of Google's Bard into Gemini and the release of its advanced version, Gemini Advanced.
- OpenAI's developments in AI agents aimed at automating complex tasks.
---
Key Topics Discussed
- Google's Transition from Bard to Gemini
- Rebranding: Google has officially transitioned its AI product from Bard to Gemini.
- Gemini Advanced: This version utilizes Gemini Ultra 1.0, reportedly reaching capabilities similar to GPT-4.
- User Experience:
- Improved conversational context (32K).
- Enhanced ability to roleplay complex scenarios.
- New features include DoubleCheck, allowing users to verify information received.
- OpenAI’s Development of AI Agents
- AI Agent Concept: Aims to evolve from simple Q&A systems to agents that can automate tasks and operate devices.
- Types of Agents:
- Device-based agents: Take control of user devices to perform tasks (e.g., transferring data, filling out reports).
- Web-based agents: Focus on tasks that can be performed online.
- Competitive Landscape: OpenAI's developments may position it against Microsoft’s Copilot.
- User Concerns About AI in the Workplace
- Survey Insights: A study from Rutgers University reveals that while many fear job losses to AI, the predominant worry relates to AI's influence over hiring and firing decisions.
- Worker Sentiment:
- 30% express concern about job elimination.
- 70% are worried about AI's role in human resource decisions.
- Apple’s Emergence in AI
- New Tools: Apple is increasing its AI capabilities, releasing an open-source model called MLLM Guided Image Editing (MGIE), which enables users to edit images using natural language prompts.
- Gemini vs. GPT-4
- First Impressions:
- Mixed feedback from early testers, with some praising its features and others noting deficiencies compared to GPT-4.
- Ethan Mollick’s Perspective: Acknowledges Gemini Advanced as comparable to GPT-4 but not superior, suggesting that both models exhibit strengths and weaknesses.
- Future Implications
- AI Agent Evolution: There is optimism that Gemini marks the beginning of a new wave in AI development towards more capable AI agents that can assist users effectively.
- Industry Shift: The emergence of another GPT-4 class model signifies a shift in the AI landscape, moving from a single-dominant model to a more competitive environment.
---
Key Takeaways
- Google’s rebranding efforts with Gemini signal a major shift in its AI strategy, aiming to compete more effectively in a rapidly evolving market.
- OpenAI’s push towards AI agents introduces new complexities and possibilities in how users interact with technology.
- Worker concerns about AI highlight the need for transparency and ethics in implementing AI solutions in the workplace.
- The competition between AI models (Gemini and GPT-4) establishes a new benchmark for capabilities and functionalities in AI development.
---
Conclusion The evolution of AI continues to unfold rapidly, with significant developments from Google and OpenAI reshaping the landscape. As these technologies advance, both their capabilities and the implications for users and workers will be critical areas to watch.
For more information, subscribe to The AI Breakdown [newsletter](https://theaibreakdown.beehiiv.com/subscribe) or follow on [YouTube](https://www.youtube.com/@TheAIBreakdown).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Today on the AI Breakdown, Google has officially changed Bar to Gemini and released the most advanced version of their Gemini model. Before that on the brief, OpenAI is working towards some seriously advanced agents. The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Go to breakdown.network for more information about our YouTube, our Discord, and our newsletter. Welcome back to the AI Breakdown Brief, all the AI headline news you need in around five minutes. As you might have guessed, our main story today is going to be about Google changing the Bard brand to Gemini, and giving access to their most advanced Ultra model, which just kicked off today.
0:40However, the information has some really interesting reporting from behind the scenes in OpenAI in a piece that they titled, OpenAI Shifts AI Battleground to Software that Operates Devices and Automates Tasks. Now, if you were an AI Breakdown listener last year, you heard me talk a lot about AI agents. One of the huge themes among developers was trying to move to an era where instead of just answering people's questions, AI agents actually had the capacity to solve their problems. In other words, you could give them a problem, and the agent could figure out what tasks were needed to solve that problem or accomplish that goal, including potentially leveraging other agents to do so.
1:20Now, this is a reality that is not here yet. There are lots and lots of companies trying and experimenting and building towards that, some of which are getting increasingly high capacity in some specific functions, but there is not yet a generalist AI agent, winner, or even really leader. OpenAI appears to be determined to change that. Writes the information, OpenAI is developing a form of agent software to automate complex tasks by effectively taking over a customer's device. The customer could then ask the ChatGPT agent to transfer data from a document to a spreadsheet for analysis, for instance, or to automatically fill out expense reports and enter them into accounting software.
1:55These kinds of requests would trigger the agent to perform the clicks, cursor movements, text typing, and other actions humans take as they work with different apps. Now, apparently they are working on actually two different types of AI agents. The one that we just described, which would take over a person's specific device, and another which would be specifically for web-based tasks. Now, this makes sense given that Sam Altman has discussed his vision for ChatGPT as ultimately a quote, super smart personal assistant for work. However, as the information also points out, it could bring OpenAI increasingly in competition with Microsoft, who are of course positioning Copilot as exactly this sort of thing.
2:29Although how far Microsoft is in any sort of plans or developments around AI agent sort of behavior isn't clear at all. There are also real questions about whether users will be comfortable with this. Right now, the only types of software that take over people's computers are malware and viruses, and so getting over that impression could be really difficult. Now, these are not new efforts, apparently. It appears that they've actually been in development for more than a year. However, there are some indications that employees within OpenAI think that these tools are going to be a really, really big deal.
2:58One of the people who was the information sources, for example, pointed out a tweet from Ben Newhouse, who was an OpenAI employee who this source said had worked on computer-using agents, and Ben on Twitter posted, building what I think could be an industry-defining zero-to-one product that leverages the latest and greatest from our upcoming models. Adding even more hype to that cryptic announcement, Pete Wellender, OpenAI's vice president of product, added, this product that Ben was describing will, quote, change everything. Now, of course, there are other indicators. These are certainly the lines that people are thinking about.
3:28In many ways, if you go back and look at how OpenAI framed custom GPTs, it was as a very first step towards an agent-like future. They, of course, also launched at their DevDay event the Assistance API, which is explicitly about helping developers build light agent-type experiences in their applications. So right now, this is just a behind-the-scenes report. There's no indication that anything is coming soon, but it's consistent with other things that we've heard, and it certainly seems to suggest where the AI arms race could be headed next. It is certainly something that, if nothing else, I will be watching closely and letting you know if I hear anything more, or frankly, probably even less definitive.
4:06Now, another company that we've been guessing at their AI strategy, and who are finally starting to tease it themselves as of Tim Cook talking to investors recently, is, of course, Apple. One of the indicators that Apple is getting deeper and deeper into its own AI strategy is the fact that they have been increasing their open source releases in the space. The latest is something called MLLM Guided Image Editing, or MGIE. It's a model that lets users use plain language to edit a photo without any photo editing software. So think of the sort of in-painting that got people so excited in MidJourney and Adobe image applications last year.
4:40Want to be wearing a different color shirt in a photo? Just say, I want my shirt to be a different color. Now, according to The Verge, the model blends two different uses of multimodal language prompts. First, it learns how to interpret user prompts, then it imagines what the edit would look like. Now, part of what makes this different is there is reasoning involved. So for example, if you had a picture of a pepperoni pizza, and you typed in the prompt to make it more healthy, MGIE would add vegetable toppings. This is very different than having to use a prompt to add vegetable toppings. Now, if you are interested in trying this out, you can download it from GitHub, or you can do a web demo over on Hugging Face.
5:16Lastly today, another interesting survey about U.S. worker attitudes towards artificial intelligence. This one comes from Rutger University's Heldrick Center for Workforce Development, which has multi-decade experience of surveying Americans around the impact of new technology in the workplace. Now, one thing that's perhaps not surprising is that there is meaningful concern among people about their job being eliminated by AI. Three in ten have that worry. However, a far more dominant worry, at least right now, has to do with the quote hidden hand of AI being involved in human resource decision making, i.e.
5:49hiring and firing. When it comes to those issues for AI, 7 in 10 US workers say they're very or somewhat concerned. As the center's director Carl Van Horn puts it, a concern about the hidden hand out there, that I'm not going to get a chance to really discuss my virtues with the hiring officer or with my boss. Instead, there'll be some algorithm that tells me whether I stay or go. So some pretty interesting nuance we're getting into when it comes to people's perceptions and fears around AI. Always interesting to see these new stats, although of course take them as one tiny piece of evidence in a much larger world.
6:22For now though, that is going to do it for today's AI Breakdown Brief. I'll be back soon with the main AI breakdown. Welcome back to the AI Breakdown. Earlier this week, an Android developer found a changelog message that suggested that we would be getting the most advanced model of Gemini this week. and that along with it, Google would be making a big marketing change, moving the Bard brand to Gemini. Now this is something that we've seen as sort of a trend among these big companies. They start with one brand, and ultimately start to settle on a cross-cutting brand that refers to everything that they're touching with AI.
6:56I told the whole story of Bing shifting to Copilot yesterday, for example, which is of course encapsulated by their Super Bowl ad, which was just released. Well, like I said at the top, the rumors are true. Bard is now Gemini and Ultra. Their most advanced model is actually available. Jack Krozek from the Google Bard team says, Today, Bard becomes Gemini. Available on web and mobile, new app in the Play Store, starting to roll out today. And introducing Gemini Advanced, access to our most capable model, Ultra 1.0. Jack continues, Bard was built to be the direct way to access Google's AI models.
7:29Last week, Gemini Pro went worldwide and completed the transition into the Gemini era. Gemini is more than state-of-the-art models. It's an ecosystem you will see through our products and APIs. Hence, Bard is now Gemini. Gemini Advanced provides access to our most capable model, Ultra 1.0. We worked with 100-plus AI expert trusted testers across multiple disciplines. They've told us they prefer Gemini Advanced for its longer context conversations, 32K, and the ability to roleplay complex scenarios. It also doesn't interrupt your flow with low rate limits. Now, from there, he goes on to a lot of other details, including discussing the new app experience on Android and iOS in the Google app, as well as another new feature called DoubleCheck, which allows users to double-check the information that's coming back from Gemini, and finally discusses the pricing structure for this most advanced Ultra 1.0 model.
8:16Users can try it for two months for free, and then it is$20 a month after. Well,$19.99 technically. Now, there are some things that are not there yet that make it not as feature-complete as ChatGPT. These include multimodal upgrades, interactive coding, deeper data analysis, file uploads, multilingual, and more. Now, there is a lot that makes this interesting. As The Verge points out,
8:47The Verge also points out the stakes. They write,
8:58of other powerful AI competitors on the market. In our test just after the Gemini launched last year, the Gemini-powered BARD was very good, nearly on par with GPT-4, but it was significantly slower. Now Google needs to prove it can keep up with the industry as it looks to both build a compelling consumer product and try to convince developers to build on Gemini and not with OpenAI. Only a few times in Google's history has it seemed like the entire company was betting on a single thing. Once that turned into Google +, and we know how that went. But this time, it appears Google is fully committed to being an AI company, and that means Gemini might be just as big as Google.
9:32Now, the only thing that I disagree with from this Verge analysis is the idea that going all in on Gemini raises the stakes for the company's ability to compete. Those stakes were raised by the very existence of leaders in the AI space that weren't Google. For a company that has been at the very forefront of innovation in this space, it was shocking last year to see how far it was behind throughout the entire year. It was in fact to many people stunning. Indeed, it put them in a position where they really had to announce Gemini in December even though they couldn't make Ultra available at that time.
10:04The exciting thing is, of course, is that because we didn't have access to the Ultra version of the model back in December, we simply had to take their assertion that it matched or exceeded GPT-4 in numerous areas at face value or choose not to believe it. Now we get to test it for ourselves, but there are some people who have had a little bit longer of a chance to already dig into it. Popular creator Marques Brownlee says, Okay, so I've been testing out Google's Gemini on a few phones for a few weeks now. Some things I've noticed that stood out. Upsides? Tons of useful new generative features.
10:34Can write letters, craft trip plans, create images, etc. All the good stuff. Notably better semantic understanding of random fact-based questions. Downside? It's missing some classic Google Assistant features like home control and adding to shopping lists. The new pop-up UI is a little more complex, but you can get used to it. Renaming it is going to confuse a lot of people for no reason. Bindu Ready from Abacus writes, My initial thoughts still continues to be somewhat nerfed and refuses to answer questions. Refuse to generate a simple illustration of George Clooney, ChatGPT is better. Missing PDF upload.
11:03Answers do seem better than the previous version. Seems to have a reasoning vibe. However, it does not answer some hard questions that GBT does. For example, it didn't get, In a room I have only three sisters. Anna is reading a book. Alice is playing a match of chess. What's the third sister Amanda doing? The answer is the third sister is playing chess. GPT-4 nails it. Overall, we plan to do a lot more analysis, but first impressions are good, but not great. TLDR, I don't think it will make a material difference to how Bard was doing before, especially if their plan is to charge for this. However, it's always good to have more players in the market.
11:34Now, someone who has a much more positive take is V. Mausiewicz, who writes, As someone who had early access, I can say that Gemini Ultra is damn impressive. When it is good, it is excellent, and this includes the most common queries, especially learning and looking up facts. I've switched to it as my default LLM. Now, Jvi is a prolific blogger and super compelling thinker, so I am inclined to take his take on this with a little bit more confidence than some of the others. And then, of course, there's Professor Ethan Mollick from Wharton. Ethan has had access to Gemini for the last six weeks and wrote an extensive post on his One Useful Thing blog giving his perspective.
12:09The post is called Google's Gemini Advanced, Tasting Notes and Implications. Subtitle, and then there were two. Now, one thing that Ethan makes clear is that he is not trying to test Gemini on the basis of benchmarks. Instead, he wanted to give a subjective mix of opinions based on his usage. He writes, let me start with the headline. Gemini Advanced is clearly a GPT-4 class model. The statistics show this, but so does a month of our informal testing. And this is a big deal, because OpenAI's GPT-4, the paid version of ChatGPT and Microsoft Copilot, has been the dominant AI for well over a year, and no other model has come particularly close.
12:46Prior to Gemini, we had only one advanced AI model to look at, and it is hard drawing conclusions with a data set of one. Now there are two, and we can learn a few things. Now I just want to stop here and put a fine point on this. It really is remarkable that for an entire year in this insanely fast-moving space, there was nothing that could really come close to GPT-4. I've actually argued in the past that having another GPT-4 class model on the scene and actually available represents a major transitional moment for the industry, from the period that was entirely defined by ChatGPT from November 2022 when it launched to basically today, to this new era that's coming, whatever it happens to look like.
13:28Now that said, Ethan continues that Gemini Advance does not obviously blow away GPT-4. He writes, it is really good, but I would concur with the test that suggests it is roughly equivalent. When it comes to the various strengths and weaknesses of these platforms, he writes, GPT-4 is much more sophisticated about using code and accomplishes a number of hard verbal tasks better. Gemini is better at explanations and does a great job integrating images and search. Both are weird and inconsistent and hallucinate more than you would like. And that gets to a really interesting section of the piece called It's Full of Ghosts.
13:57Ethan writes, no one has a great definition of sentience, which is okay because LLMs are in no way sentient. They are software systems designed to create human-like language. But there is a weirdness to GPT-4 that isn't sentience, but also isn't like talking to a program. A weirdness that only comes out after you spend enough hours playing with the AI and getting unnerved or delighted or both by its unexpected abilities and seeming intelligence. There was a famous controversial paper put out by Microsoft Research soon after the release of GPT-4 called Sparks of Artificial General Intelligence that tried to put this argument into scientific terms, but ended up just calling it sparks of artificial general intelligence.
14:31It is the illusion of a person on the other end of the line, even though there is nobody there. GPT-4 is full of ghosts. Gemini is also full of ghosts. Seriously, if you use the system for a while, I can almost guarantee at least one moment when you stand up from your desk, walk around the room, and wonder what is going on. Now his takeaway is that the sparks that we saw in GPT-4 are not based on GPT-4, but are byproducts of GPT-4 class models. From a tone and personality perspective, he suggests that GPT-4 is more bland, where Gemini is more friendly, agreeable, and has a, quote, tendency towards wordplay.
15:05But ultimately, he says these models are really, really similar. The other thing that he argues about Gemini is that it, quote, illuminates a vision of AI as a powerful integrated personal assistant. He basically argues that the barred integration with the Google ecosystem of Gmail, Google Docs, travel tools, etc. was interesting but too dumb to actually use. Whereas now, he says, with a smarter brain in the form of Gemini Advanced, you can start to do some really interesting things that, at their best, seem magical. Go through my emails, tell me which are important, and draft replies for each.
15:33Look up my next conference and plan a trip I would like. He says it isn't there yet, but it is very much closer to being an actual assistant rather than the limited series at Alexis we have seen in the past. That is, in part, why I suspect that Gemini Advanced is the start, not the end, of a wave of AI development. We can start to see a world where AI agents act on our behalf. A GPT-4 class model is not quite strong enough to power these agents, but we are getting close. So really interesting stuff. And ultimately what we come back to is that we are now living in a two GPT-4 level world, which I think is going to have some fairly significant implications that we will discover as it happens.
16:09For now though, an exciting day in the AI world. I appreciate you listening or watching as always. And until next time, peace.
16:20Thank you.
From the publisher
Bard is no more! Bard has become Gemini and Gemini now features Gemini Advanced, which uses Gemini Ultra 1.0 -- the first non OpenAI model to hit GPT-4 levels. Reports also suggest that OpenAI's next big play is AI agents.
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI.
Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe
Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown
Join the community: bit.ly/aibreakdown
Learn more: http://breakdown.network/
