In short
The AI Daily Brief: Episode Summary
Podcast Title
The AI Daily Brief (Formerly The AI Breakdown)
Episode Title
Meta Launches Midjourney, Dall-E-3 Competitor
Episode Description
This episode discusses Meta's new image generation tool, Imagine, and its comparison to existing tools like Midjourney and DALL-E 3. The episode also reflects on Google's recent announcement of Gemini and its implications for AI development.
---
Key Topics Discussed
- Google Gemini Announcement
- Initial excitement about Gemini being a competitor to GPT-4.
- Critiques arose regarding the authenticity of its demo videos:
- A widely circulated demo of Gemini identifying drawings in real-time was revealed to be staged.
- The human-drawing interaction was not as described, leading to skepticism about the model's capabilities.
- Despite the criticisms, Google's stock rose by 5.3% after the announcement, suggesting investor optimism.
- Market Reactions and AI Developments
- Analysts noted that Google's Gemini could address user concerns related to OpenAI's recent updates.
- Amazon's CEO emphasized the transformative impact of generative AI on customer experiences, hinting at significant developments in their AI technology.
- AI in Fast Food Industry
- McDonald's announced plans to integrate Google AI into their operations to enhance service efficiency and food quality.
- Concerns emerged regarding AI's potential role in replacing human workers, although some studies indicate that human oversight remains critical in the AI deployment process.
- Shareholder Activism in AI Governance
- Chris Novoselic, former bassist of Nirvana, called for Microsoft to assess the impact of its AI products more responsibly, arguing against rapid deployment without proper safety measures.
- This raises questions about the balance between innovation and ethical considerations in AI development.
- Geopolitical Implications of AI
- G42, a UAE AI company, announced it would phase out Chinese hardware to comply with U.S. regulations, highlighting the geopolitical tensions surrounding AI technology access.
- Introducing Meta's AI Image Generator: Imagine
- Meta announced over 20 new generative AI tools aimed at enhancing user experience on its platforms.
- The Imagine tool has been developed as an alternative to existing image generators like Midjourney and DALL-E 3.
- Mixed results reported regarding its performance, particularly in generating realistic human figures.
- Integration of AI Across Meta Platforms
- Meta aims to integrate generative AI tools into existing applications like Facebook and Instagram for improved user experiences.
- Features include a virtual assistant with enhanced capabilities and an ability to create images within chats.
- Safety and Ethical Considerations
- Meta introduced invisible watermarking for AI-generated images to address copyright concerns.
- Commitment to safety through processes like multi-round automatic red teaming (MART).
---
Key Takeaways
- Market Positioning: Meta's recent launch appears to be a strategic response to Google's Gemini and aims to solidify its position within the competitive AI landscape.
- User Experience Focus: The emphasis on integrating AI into everyday applications suggests a shift towards making AI tools more accessible and user-friendly, rather than solely focusing on groundbreaking AI capabilities.
- Ethical Responsibility: Shareholder activism and new safety measures indicate a growing awareness of the ethical implications of AI deployment.
- Community Engagement: The AI Daily Brief's promotion of an education and learning community highlights the importance of user engagement in the evolving AI landscape.
---
Conclusion In this episode of The AI Daily Brief, discussions centered around the competitive dynamics of AI technology, particularly focusing on Meta's new offerings and the implications of Google's recent developments. The conversation emphasized the significance of user experience, ethical considerations, and the integration of AI into existing platforms.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:01Today on the AI Breakdown, we're talking about Meta's new AI announcements, including their Imagine Image Generator. Before that on the brief, more thoughts on Google Gemini. The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Go to breakdown.network for more information about our YouTube, our Discord, and our newsletter.
0:24Welcome back to the AI Breakdown Brief, all the AI headline news you need in around five minutes. Well, friends, obviously the big story from this week was the surprise announcement of Google Gemini. We discussed in that immediate reaction video a little bit about how, while initially everyone's reaction was to be impressed that there was finally a GPT-4 level competitor, pretty quickly there started to be some questions. The much-touted MMLU results that showed Gemini not only beating GPT-4, but also beating human expertise, weren't really carried out in the same way that tests including GPT-4 were carried out, which was a little bit of a wet blanket on that announcement.
1:02And then on Thursday, a lot of the chatter on Twitter and in other parts of the internet was about how one of the most shared demo videos from the launch, which showed a company representative speaking to Gemini as it identified drawings in real time, was basically entirely edited and wasn't really at all what it seemed. The video that I'm referring to is of course the one where a human is drawing something that starts off squiggly lines and then becomes a duck, and Gemini seems to respond in kind, explaining what's going on, and then it moves to inventing a game, then it solves visual puzzles.
1:33Well, basically it didn't happen as it was presented, which made it seem like the person was just speaking to Gemini right there. Instead, as Bloomberg opinion columnist Parmy Olson put it, Google filmed those hands doing a bunch of things and then showed stills of the footage to Gemini one by one. There was no voice conversation but a text exchange like you'd have with ChatGPT or Bard. Google's video made it look like you could show different things to Gemini Ultra in real time and talk to it. You can't. The voice you hear in the video is just reading the text prompts. So again, all of this just served to throw a little cold water on this, but ultimately, the bigger point still remains that we just don't have Ultra yet, and until we do, there are going to be major questions.
2:11Yet at the same time, even though Google's stock price was initially flat on the announcements, Thursday saw a significant uptick as Wall Street digested the news. Alphabet shares ended 5.3 % higher on the day, with analysts believing that the new Gemini model could narrow the gap with OpenAI and, by extension, Microsoft. JP Morgan analysts wrote, Google is beginning to address investor concerns around generative AI and the high cost of running GenAI models through the combination of Gemini's different model sizes. Other analysts noted that the introduction at Gemini comes at a time when some people are more skeptical of OpenAI's models.
2:46Wrote McCary analysts, Gemini's release comes at an interesting point in time where OpenAI and ChatGPT users have been complaining about how updates to the GPT model family have potentially impacted the quality of its output. If Google is shipping a GPT-4 beating model, this could help gather user and developer momentum behind Google. Now, staying in big tech land for a moment, Amazon CEO Andy Jassy didn't announce anything new, but in an interview with Jim Cramer, he did say that generative AI will, quote, change every customer experience. He said, if you've studied generative AI and you're still scoffing, you're really not paying attention.
3:21Generative AI is going to change every customer experience, and it's going to make it much more accessible for everyday developers and even business users to use. So I think there's going to be a lot of societal good. He also liked Amazon's prospects in the AI race, saying, We think we have a real opportunity to be a leader there, and we're in the process of building a much more expansive large language model underneath Alexa that will make her both much more knowledgeable and much more conversational. Is this the large Olympus model that has been rumored? We will just have to wait and see. Now, of course, it's not just the big tech companies that are exploring how to use AI, but the CPG companies and the restaurants.
3:55For example, The Verge writes, McDonald's will use Google AI to make sure your fries are fresh or something? The fast food company says it will be applying generative AI to its operations starting in 2024. The TLDR, McDonald's is partnering with Google to bring generative AI to, quote, thousands of stores through hardware and software upgrades. That includes ordering kiosks, the company's mobile app, and data science that helps them optimize operations. The company claims one of the outcomes of that will be hotter, fresher food, to which The Verge says, It's not completely clear what that means, but we can read between the lines.
4:27Expect more AI-driven automation at a drive-thru near you in the coming years. The Verge also notes, quote, The McDonald's statement skirts the question of AI replacing human workers, mentioning only that the system should reduce complexity for store crews and that it will, quote, power exciting new experiences for crews and customers. There's a whiff of robots replacing human workers in the air, but some other research suggests that while very real, the value proposition of AI for drive-thru is actually driven by real people. Writes Bloomberg, Checkers and Carl's Jr. are among U.S. fast food chains hailing AI-powered drive-thrus as labor-zapping wizards that speed up service.
5:02But a popular provider of these systems recently revealed a crucial part of how it gets so many orders right. Humans. Basically, this company, Presto Automation, has a chatbot that's meant to take orders with very little human intervention. However, disclosures from recent SEC filings say that, quote, off-site agents working in locales such as the Philippines help during more than 70 % of customer interactions to make sure AI systems don't mess up. This is a humans-in-the-loop approach to AI that I think we're going to see a lot more as companies figure out how to actually bring artificial intelligence into their workflows.
5:33Here's a wild one for a guy like me whose first AOL screen name in 1995 was Kurt C. Freak. The original bassist and founder of Nirvana, Chris Novoselic, recently served as the spokesperson for a Microsoft shareholder initiative asking the Redmond giant to study and report on the impact of its AI initiatives. Novoselic, on behalf of these shareholders, argues that Microsoft has rushed generative AI products to market and wasn't giving enough consideration to guardrails. In a recorded video message, he said, When Microsoft released its generative AI-powered Bing last February, numerous AI experts and investors expressed concern.
6:08Many urged Microsoft to pause and consider all the risks associated with this new technology so that the company could establish risk mitigation practices. Yet our company raced forward releasing this nascent technology without the appropriate guardrails. Novoselic has apparently been a long-term Microsoft shareholder and added, Generative AI is a game-changer there's no question, but the rush to market seemingly prioritizes short-term profits over long-term success. Now, outside of the weird pop culture connection of this, I think it's worth pointing out that if one is really interested in slowing down the AI arms race, shareholder activism, i.e.
6:41internal pressure, that shifts the balance of market pressure, could be a very promising vehicle, especially compared to people just lobbing complaints in from the outside. Moving to the geopolitics of AI for a moment, an update in a story that we've been covering, which is the United Arab Emirates company G42. We recently had that big New York Times feature about them and speculated a little bit about the nature of their relationship with OpenAI, which was just announced in October. And part of the tension was that they were deeply enmeshed with both China and the US. However, it appears that they are no longer able to be any sort of neutral middle.
7:14And the Financial Times reports, quote, UAE's top AI group vows to phase out Chinese hardware to appease U.S. Abu Dhabi-backed G42 says it cannot work with both sides and retain access to American-made AI chips. Said Chief Executive Peng Shao, for better or worse as a commercial company, we are in a position where we have to make a choice. We cannot work with both sides. We can't. Now, depending on your take, this could be an example of U.S. policy and pressure actually working, at least in terms of its stated objectives of denying China access to advanced AI infrastructure. It's a really interesting one for sure, and certainly something I'm going to continue watching.
7:51Lastly, I want to call out a viral project that has been flying all around Twitter slash X this week called Magic Animate. It's basically a model by which you can take a static image, which can be from the world or can be generated by something like Midjourney or Dali 3, reference it against a motion file from Meta's dense pose, and get an actual animation of that seed image that follows the motion from the reference file. I actually just did a tutorial of this for the AI Breakdown Learning Community beta that is happening right now. We're just coming to the close of the first week of this beta, where we're doing tutorials, sharing case studies, and participating in follow-up challenges that get people actually trying tools.
8:28Now, just before the main part of this episode, I'm going to share a little bit more about how you can get involved in January, if that's something that's interesting to you. But for now, if you're interested, go on Twitter slash X and see some very cool things indeed. However, that is going going to do it for today's AI Breakdown Brief. Next up, the main AI breakdown. Hey guys, before we get into the main part of the episode, I wanted to mention just briefly that we are now in the midst, we're actually just closing out the first week of the AI Breakdown AI Education and Learning Beta. This is a community of learners where each day I'm dropping in tutorials, case studies, challenges, and a community of people are discussing them, going out and doing those challenges.
9:07In other words, learning AI by doing, and getting a chance to ask questions and talk with people who are experiencing similar problems, taking advantage of similar opportunities, and generally adapting to this new AI-powered world. I'm incredibly encouraged by how it's going so far, and in about a week I'll be opening up registration for next month's second beta test for January. For now, I wanted to let you guys know that that was coming, and if you are interested in getting on the waitlist for that, go to bit.ly slash AI beta. You'll see the short write-up that I did of December's beta, plus a link to a form where you can sign up for the waitlist.
9:39I'd love to have you participate in January. So again, that's bit.ly slash AI beta. And now let's get to the main episode. In what I would consider highly unsurprising news, Meta has made a slew of new announcements about AI features and tools. So first, why do I say it's unsurprising? Well, there's two reasons. One, the chaos and turbulence at OpenAI has provided an opportunity for all of the players who are not OpenAI to try to remind the world that they exist and exist as alternatives. But two, especially in a week where Google has answered that opportunity by announcing their most advanced model ever in Gemini, it stands to reason that Meta would try to get a piece of that narrative action as well.
10:23So what did we get? Well, on December 6th, the company announced that they were testing more than 20 new generative AI tools. They write, to close out the year, we're testing more than 20 new ways generative AI can improve your experiences across Facebook, Instagram, Messenger, and WhatsApp, spanning search, social discovery, ads, business messaging, and more. Now, right off the bat, we see something that we've been talking about a lot on this show, which is the fact that we are moving to a period where it's not just about advanced capability announcements, but also about integration into the tools that we're already using.
10:53Right up front, they say these 20 generative AI tools are going to improve people's experiences on the platforms they're already using, Facebook, Instagram, Messenger, and WhatsApp. So what are some of these 20 things that they're testing now? Well, one is an update to Meta AI, which is their virtual assistant. They write, We're making it more helpful with more detailed responses on mobile and more accurate summaries of search results. We've even made it so you're more likely to get a helpful response to a wider range of requests. They write that Meta AI is now helping outside of chats as well.
11:21Quote, It's doing some of the heavy lifting behind the scenes to make our product experiences on Facebook and Instagram more fun and useful than ever before. The large language model technology behind Meta AI is used to give people in various English language markets, options for AI-generated post comment suggestions, and community chat topic suggestions in groups, serve search results, and even enhance product copy in shops. Basically what they're saying is that in addition to their assistant chat bot getting better, the same technology, the same LLM that underpins that bot is being used and integrated in effectively everywhere that text lives inside meta experiences, be it groups or the shopping experience or what have you.
11:56Now part of the upgrade with Meta AI is a new ability to create images inside chats. Specifically, they've added a new remix feature that they're calling Reimagine. They write, Here's how it works in group chat. Meta AI generates and shares the initial image you requested, and then your friend can press and hold on the picture to riff on it with a simple text prompt, and Meta AI will generate an entirely new image. Now you can kick images back and forth, having a laugh as you try to one-up each other with increasingly wild ideas. Basically, this is a social extension of an image generation tool.
12:24Another new upgrade to the Meta AI Assistant is that it can pull in reels, which could be useful for things like planning trips. So you can, for example, see reels of the places that you might want to visit. And they say, this is just the first example of how we'll build even deeper integrations across our apps to make meta AI an even more connected and personal assistant over time. And you'll see at this point already that so many of these things that they're experimenting with are not big newsmaking changes, but just different workflows, user experience integrations and updates, and basically bringing AI tools effectively everywhere in the meta suite of apps.
12:55Like I said, anywhere that there's language, they're testing recommendations. For example, they talk about how they're helping creators respond to their fans. We want to give creators generative AI tools to help them work more efficiently and connect with more of their community. We're starting to test suggested replies and DMs to help creators engage with their audiences faster and more easily. Now, for those of you who have been experimenting with the character AIs, no relation of course to character.ai, that Facebook slash Meta announced earlier this year, the big update here is that for a number of them, they've added long-term memory.
13:24Basically, they allow you to go away from a conversation and pick it back up right where you left off. Now, of course, requisitely, there also is an update around responsibility and safety. One of the banner announcements is that they're adding a new watermark to any images that are created with Meta AI's tools. They write, while it's imperceptible to the human eye, the invisible watermark can be detected with a corresponding model. It's resilient to common image manipulations like cropping, color change, screenshots, and more. We aim to bring invisible watermarking to many of our products with AI-generated images in the future.
13:53Now, more broadly, when it comes to safety, they write, we're continuing to invest in red teaming, which has been a part of our culture for years. As a part of that work, we pressure test our generative AI research and features that use large language models with prompts we expect could generate risky outputs. Recently, we introduced multi-round automatic red teaming, or MART, a framework for improving LLM safety that trains an adversarial and target LLM through automatic iterative adversarial red teaming. We're working on incorporating the MART framework into our AIs to continuously red team and improve safety.
14:21So like I said, big theme here is not crazy banner new. We've got Llama 3 style announcements as hungry as some people are for that. Instead, it's clearly all about the integration phase of AI and how Meta's existing tools are going to find their way to be actually useful across their family of experiences. That said, to the extent that there was something that captured a lot of people's attention, it's that Meta has pulled its AI image generation out of the places that it lived inside its apps, and created a public space for it at imagine.meta.com. So like Midjourney or Stable Diffusion or Dolly 3, this is an image generation model, although this one has been trained on Facebook and Instagram photos.
15:01We'll see if that makes a difference in terms of what it's good at in just a moment. As I said, it exists outside of the apps at imagine.meta.com, but it does require a meta account to sign in with. Results so far are somewhat mixed. VentureBeat writes, VentureBeat's brief, unscientific tests showed that it only sporadically produced realistic human figures and structures. Often our imagery included strange glitches like melted body parts and scenery. They also point out that a lot of the features that we've gotten used to with other tools like MidJourney just don't exist. There's no option to remix from Imagine.Meta.com, although you can do that in its Messenger apps.
15:32There's also no way to resize images. And on top of that, of course, there is controversy around how the model was trained. Now this is built on top of Meta's own AI model that's called EMU, but the issue is, of course, that it was trained on people's Facebook and Instagram images, although excluding things that were shared in private messages. Still, that hasn't stopped some people from saying that this is not how they ever intended for their images to be used, and that this represents a problem. Carla Ortiz writes, Outrageous! Meta released a generative AI model trained on 1.1 billion images from Facebook and Instagram.
16:01Copyrighted works, our pictures and our loved ones' pictures all used to train this model. Tech companies are claiming ownership of things they do not own. This must stop. That said, a lot of people also think that this is actually probably Facebook trying to avoid copyright violation issues because its terms of service for users who share photos to its site are probably a lot more accommodative of this type of work than just scraping artwork from the public web. Now, it seems to me, and this is the impression that I think a lot of people have, that this initial version of Imagine is really meant to be a competitor to free image generation tools.
16:31It's not meant to stand next to MidJourney's paid model or anything like that, and so that may be a useful way of judging it. It's got a super simple user interface, which would suggest for that as well. The couple of tests I did weren't that bad. I did rugged Santa, cinematic shot, bokeh, Christmas tree farm, misty background, and got some decent results. Blaine Brown writes, First set of images via Meta's emu slash imagine generator is impressive. Prompt, a baby snow dragon next to a pine tree in a winter wonderland. Cinematic. You can see, and this is something that people have noted, that it handles detail pretty well.
17:01Chase Lean, who often does comparisons between different image generators, did his own test to compare it to Mid-Journey Dolly 3 and Adobe Firefly across 10 image categories. One was a realistic photo, two were landscape photos, and by the way, if you're watching this, it is the upper left that is meta-imagine each time, product photos, text generation, which spoiler alert, it wasn't really able to do, vector graphics, pixel art, which meta didn't do at all, interior designs, close-up shots, wildlife photos, and coloring pages for a kid's book. Chase's summary, meta's AI is good at making realistic photos.
17:34This is probably due to their training data. It's a decent free option and slightly better than Dolly 3 in this regard. Dolly 3 usually makes more cartoonish picks. However, it is clearly outmatched by Adobe Firefly and MidJourney, though both are paid. Dolly 3 also beats it at text generation and the ability to follow instructions inside the prompt. So, is Imagine Revolutionary? Absolutely not. And in fact, I actually think it points to something different that's happening that is an important trend, which is the commoditization of advanced AI models. This year, things are still happening so fast that there are meaningful differences between how good these various models are.
18:08It's not just image generators, but also text generators and LLMs as well. However, at what point do things consolidate around a version that's so good that even if there are advances being made with GPT-5 and GPT-6, the vast majority of the world is using, for example, Dolly 3 slash MidJourney 5.2 level quality image generation and GPT-4 level text generation. It feels like that's happening fast, and in that world, competition is going to be much more about exactly what Meta seems to be trying to compete on, which is integration into other apps and experiences. So even if you are not overly impressed with the quality of images just yet, I still think it's worth keeping an eye on what they're doing, given how much access to people's existing behavior they have, and how easily they can bring these tools into those experiences.
18:52However, for now, that's going to do it for today's AI Breakdown. I appreciate you listening or watching as always. Until next time, peace.
From the publisher
How does the new tool stand up to other image generators? And what else did Meta AI announce. The day following Google's Gemini was a mixed bag for the company. On the one hand, socials acted betrayed on discovering the most impressive demo video had sort of been staged. On the other hand, Wall Street drove the stock up 5+% - adding $80B to the company's market cap.
Interested in the January AI Education Beta program?
Learn more and sign up for the waitlist here - https://bit.ly/aibeta
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI.
Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe
Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown
Join the community: bit.ly/aibreakdown
Learn more: http://breakdown.network/
