ChatGPT Can Now See and Hear

25 Sep 2023 · 17 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The AI Daily Brief: Episode Notes

Podcast Title

The AI Daily Brief Formerly The AI Breakdown: A daily news analysis show focused on artificial intelligence.

Episode Title

ChatGPT Can Now See and Hear Episode Description: Discussion of new multimodal features in ChatGPT and Amazon's investment in Anthropic.

---

Key Highlights

  1. Introduction
  2. The episode covers major updates regarding ChatGPT's capabilities and a significant investment by Amazon in Anthropic.
  3. The format includes daily news and discussions around important AI developments.
  1. Amazon Invests in Anthropic
  2. Investment Amount: Amazon will invest up to $4 billion in Anthropic, an AI competitor.
  3. Anthropic's Differentiation:
  4. Notable for its chatbot Claude and a 100k context window.
  5. Focuses on AI safety through "constitutional AI," promoting principles-based reasoning rather than reinforcement learning.

Financial and Collaborative Aspects

  • Amazon will take a minority stake, ensuring Anthropic's governance remains unchanged.
  • Collaboration includes:
  • AWS becoming Anthropic's primary cloud provider.
  • Development of future AI models utilizing Amazon’s technology (Tranium and Inferchia chips).
  • Support for Amazon’s Bedrock platform to optimize enterprise model functionalities.
  1. Meta's New AI Chatbots
  2. Meta is reportedly creating AI chatbots with distinct personalities to engage younger users.
  3. Challenges arise as TikTok overtakes Instagram in popularity among teens.
  1. Microsoft's Engagement in Nuclear Energy
  2. Microsoft is exploring nuclear energy to power data centers, potentially employing small modular reactors (SMRs).
  1. China's Chip Factory Initiative
  2. China plans to bypass U.S. chip sanctions by constructing a chip factory powered by a particle accelerator, highlighting the influence of U.S. technology policies.
  1. Proposed U.S. Executive Order on AI
  2. The White House may introduce an executive order for cloud companies to disclose AI usage, aiming to monitor potentially risky AI developments.
  3. This could treat computing power as a national resource.

---

Major Developments in ChatGPT

  1. New Multimodal Features
  2. Capabilities: ChatGPT can now process audio and images, enhancing human-computer interaction.
  3. Examples:
  4. Users can send images for assistance (e.g., lowering a bike seat).
  5. Users can carry on voice conversations with ChatGPT, asking for stories or advice.
  1. Technical Details
  2. Speech Recognition: Utilizes Whisper for transcribing spoken input.
  3. Voice Generation: A new model creates human-like voices, developed with professional voice actors.
  4. Deployment: New capabilities are rolled out gradually to ensure quality and safety.
  1. Strategic Context
  2. Competitive pressure from Google's upcoming Gemini pushes OpenAI to adopt these new multimodal features quickly.
  3. The advancements position ChatGPT as a more powerful assistant, potentially competing with Google searches for everyday tasks.

---

Speculative Insights

  1. Internal Development at OpenAI
  2. Discussions on Reddit hint at a new model named "Iraqi" that could handle multimodal inputs (text, image, audio, video) with reduced hallucination rates.
  3. Speculation about another model, "Gobi," which may represent significant advancements toward AGI (Artificial General Intelligence).
  1. Twitter Conversations on AGI
  2. Various Twitter accounts are discussing potential breakthroughs in AGI, raising the stakes for AI technology's future.
  3. Concerns arise regarding the implications of rapid advancements and the potential risks associated with AGI development.

---

Conclusion

  • The episode emphasizes the rapid evolution of AI technologies, particularly multimodal capabilities in ChatGPT.
  • The competitive landscape is intensifying, with significant investments and innovations shaping the future of artificial intelligence.
  • Speculation surrounding AGI development adds a layer of urgency and intrigue to future discussions.

---

Subscribe for Updates

  • [AI Breakdown Newsletter](https://theaibreakdown.beehiiv.com/subscribe)
  • [YouTube Channel](https://www.youtube.com/@TheAIBreakdown)
  • [Community Join Link](https://bit.ly/aibreakdown)

Final Thoughts

  • The episode marks a pivotal moment in AI development, with October expected to be filled with groundbreaking updates and features.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Today on the AI Breakdown, we're looking at some massive new multimodal features from ChatGPT. Before that on the brief, Amazon makes a major investment in OpenAI competitor Anthropic. The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Go to breakdown.network for more information about our Discord, our newsletter, and our YouTube channel.

0:24Welcome back to the AI Breakdown Brief, all the AI headline news you need in around five minutes. Very, very late last night slash early this morning, news broke that Amazon was making a huge, up to$4 billion investment in Anthropic. Anthropic is of course the company best known for their chatbot Claude, and has tried to differentiate itself from OpenAI in a couple of different ways. One is from a feature standpoint. While ChatGPT remains at a much lower context window, Claude rolled out a 100k context window earlier this year, and Anthropic has also tried to differentiate itself based on its approach to AI safety.

0:58Rather than focus on reinforcement learning through human feedback, Anthropic has invested in something that they call constitutional AI. The idea is to train the AI on a set of underlying principles drawn from a variety of different sources and effectively help it reason around what it should or shouldn't do in any given situation based on those underlying principles. That idea of safety does make it into the company's press release. In fact, the title of the announcement on Anthropics webpage is Expanding Access to Safer AI with Amazon. Now, in terms of financial details, they didn't reveal much.

1:30Anthropic writes as part of the investment, Amazon will take a minority stake in Anthropic. Our corporate governance structure remains unchanged with the long-term benefit trust continuing to guide Anthropic in accordance with our responsible scaling policy. As outlined in this policy, we will conduct pre-deployment tests of new models to help us manage the risks of increasingly capable AI systems. So no valuation given here. We just know presumably that if Amazon deployed the entirety of that amount, they would still own less than 50 % of the company. Now, a lot of emphasis in the announcement is around how the companies will be working together beyond just capital.

2:03They write, The agreement is part of a broader collaboration to develop the most reliable and high-performing foundation models in the industry. Our frontier safety research and products, together with Amazon Web Services' expertise in running secure, reliable infrastructure, will make Anthropics' safe and steerable AI widely accessible to AWS customers. So specifically, AWS is becoming Anthropics' primary cloud provider, and in addition, they are committing to train future models on Amazon's AWS Tranium and Inferchia chips, with the idea being that in addition to just using these new chips, they will help develop the future versions of them as well.

2:35As part of the announcement, Anthropic also said that they're expanding support of Amazon's Bedrock platform. Bedrock is Amazon's approach to giving enterprises access to multiple models from a single space, and Anthropic writes their increased support includes secure model customization and fine-tuning on the service to enable enterprises to optimize Claude's performance with their expert knowledge while limiting the potential for harmful outcomes. Now, obviously, everyone understands instantly upon reading this that this is Amazon's Microsoft OpenAI style deal. It is Amazon going deep with one of the leading startups in the foundation model space, and it is clearly a multi-pronged partnership designed to touch on everything from Amazon's enterprise services all the way to their development of new AI chips.

3:15So you have Microsoft teaming up with OpenAI, Amazon teaming up with Anthropic, and Meta, Google, and presumably Apple all going in on their own. Speaking of Meta, we've had reports for a few weeks that the company was developing AI chatbots with different personalities in an attempt to increase engagement among younger users of their services. According to the Wall Street Journal, those chatbots could be coming as early as this week. WSJ explains the context, saying, Going after younger users has been a priority for Meta with the emergence of TikTok, which overtook Instagram and popularity among teenagers in the past couple of years.

3:45The shift prompted Meta chief executive Mark Zuckerberg in October 2021 to say the company would retool its teams to make serving young adults their North Star rather than optimizing for the larger number of older people. Now, in terms of the personalities of these bots that are actually coming, the WSJ writes about one called Bob the Robot, which is a self-described SaaS master general with superior intellect, sharp wit, and biting sarcasm. Now, the reference point is the robot bender from Futurama, but I have to say, the description of a SaaS master general makes me wonder just how in touch with the kids this company really is.

4:18Now, obviously, other social media companies like Snap have been experimenting with chatbots to increase engagement as well, with frankly inconclusive results so far. Moving back over to the world of Microsoft for just a moment, a job listing got some people chattering over the weekend. Data Center Dynamics sums up Microsoft Cloud hiring to implement global small modular reactor and micro reactor strategy to power data centers. Basically, it appears that Microsoft is expanding their engagement with nuclear energy and is potentially exploring how SMRs or small modular reactors could be a part of their energy mix in the future.

4:50Radiant Energy Fund's Mark Nelson writes, Word is out, Microsoft is plunging ahead on nuclear energy. They want a fleet of reactors powering new data centers. And now they're hiring people from the traditional nuclear industry to get it done. A world is coming where only the tech companies willing to become nuclear power developers may get to keep expanding their cloud businesses, and only countries open to new reactors get to host this expansion. A world where tech companies with 50 % margins become the only survival hope for traditional industrial concerns, with 5 % margins who need someone else to bootstrap a proper electricity supply.

5:20The race is on. Now, speaking of advanced technology powering new data centers and other manufacturing concerns, the South China Morning Post is reporting that China is planning to get around U.S. chip sanctions by building a massive chip factory powered by a particle accelerator. SEMP writes, China is exploring new avenues to bypass restrictions on lithography machines which are used in the production of microchips. Using particle accelerators to create a novel laser source, researchers are laying the foundation for the future of semiconductor fabrication. Now, what's interesting to me about this story is just the way that it reflects how much is in flux right now and how much U.S.

5:54policy towards China around artificial intelligence-related technologies is having an impact in how that company plans its technological future. Speaking of the U.S. government, Semaphore is reporting that the White House is looking into an executive order on artificial intelligence that would, among other things, force cloud companies to disclose which AI companies were using their services. From the article, the provision would direct the Commerce Department to write rules forcing cloud companies like Microsoft, Google, and Amazon to disclose when a customer purchases computing resources beyond a certain threshold.

6:23The order hasn't been finalized and the specifics of it could still change. Now, Semaphore draws the connection to KYC rules for banking, and basically this is another way for authorities to have a sense of who's making extreme transactions, although this case in energy, in order to effectively get out of challenges before they happen. As Semaphore writes, the rules are intended to create a system that would allow the U.S. government to identify potential AI threats ahead of time, particularly those coming from entities in foreign countries. If a company in the Middle East began building a powerful large language model using Amazon Web Services, for example, the reporting requirement would theoretically give American authorities an early warning about it.

6:57Now, the last piece is also really interesting. They write, the policy proposal represents a potential step towards treating computing power like a national resource. Hold aside the specifics of this potential executive order. There are a lot of people who sit at the intersection of global politics and technology who think that that idea that computing power is a national resource is a good one for the government to embrace. But for now, that is going to do it for today's AI Breakdown Brief. We are kicking off the week with a bang. Stick around. We're going to talk a little bit more about Open AI's latest multimodal announcements, plus some big speculation from Reddit, all of that and more coming up on the main episode.

7:32Hey guys, one more quick thing before we get into the main episode. If you subscribe to the newsletter, you've seen this and you might have heard it on an earlier episode, but right now I am getting information from you guys, the listeners, about what you are looking for in terms of AI educational resources. A bunch of you have filled out the survey already, and it's so helpful, but if you would take the about one minute and go to bit.ly slash AI breakdown survey, I would love to know what type of online courses you might need, what you're trying to learn more about, whether you'd be interested in a community of learners.

8:02I'm getting really close to making some decisions about what we're going to do next, and I really want all of your input. Again, it'll take about one minute and you can find it at bit.ly slash AI breakdown survey. Thanks so much. And now on with the show. Welcome back to the AI breakdown. Today, we are talking about the latest developments in chat GPT. They are very emblematic of the larger business battle that we find ourselves in the midst of. And in the second part of the show, we will dig into some very intriguing rumors coming from Twitter and Reddit around the company and just how much they've developed internally that we don't yet have visibility into.

8:37But let's kick it off with the announcement from this morning that, as they put it, ChatGPT can now see, hear, and speak. Developer relations Logan over at OpenAI says, This is one of the biggest evolutions for ChatGPT to date. Y 'all are going to love these new capabilities. Truly incredible. So there are two big things going on here. The first is around using images as inputs for ChatGPT. The example they give is they take a photo of a bike and ask ChatGPT to help them lower the bike seat. ChatGPT responds, giving them a set of instructions and then saying, if you have tools, show me and I'll guide you further.

9:10The prompter takes a close-up photo of a specific part of the bike and draws a circle to let ChatGPT know to focus on that specific part. The prompter Ryan says, is this the lever? To which ChatGPT responds, no, that's not a lever, it's a bolt. You'll need an Allen wrench to loosen it. Now, obviously, we don't need to go too deep into the details here, But the point is that all of a sudden you can use pictures of the real world to interact with ChatGPT in a way that wasn't possible before. This, of course, opens up a huge number of different use cases, which is why people have been excited about multimodal and image-based inputs.

9:41Now, the second part of it is that in addition to just using voice as an input for the ChatGPT mobile app, ChatGPT can now talk back. OpenAI writes, use your voice to engage in a back and forth conversation with ChatGPT. Speak with it on the go, request a bedtime story, or settle a dinner table debate. They've been loving the cutesy examples recently, and the bedtime story is the one that they chose to demo. Now, one small technical detail that was interesting. When it comes to speech recognition, they use Whisper, which has, of course, been lauded for being much farther ahead than many other text recognition services.

10:11And that's what's used to transcribe when someone speaks into ChatGPT. But they write, the new voice capability is powered by a new text-to-speech model, capable of generating human-like audio from just text and a few seconds of sample speech. We collaborated with professional voice actors to create each of the voices. They have five different voices, Juniper, Sky, Cove, Ember, and Breeze, that they give a demo of. So, is this all rolling out all at once? The answer, of course, is no. OpenAI writes, we are deploying image and voice capabilities gradually. They basically say that this is their normal model anyways, but when it comes to things like image and voices, it's even more important.

10:43They write, the new voice technology, capable of creating realistic synthetic voices from just a few seconds of real speech, opens doors to many creative and accessibility-focused applications. However, these capabilities also present new risks, such as the potential for malicious actors to impersonate public figures or commit fraud. When it comes to the challenges of new image inputs, they say they range from hallucinations to people overly relying on the model's interpretation of images in high-stakes domains. So TLDR, this is the update that the information was reporting about about a week ago.

11:10The context they gave, which is the one I agree with, is that the impending reality of Google's Gemini is creating pressure for OpenAI to race towards multimodality, perhaps faster than they might otherwise have. That was Dr. Jim Phan from NVIDIA's take when DALI 3 was announced. He tweeted, It's not just a stance against MidJourney, it's actually a sneak peek of the upcoming epic battle of massively multimodal LLMs against DeepMind Gemini. Now, the interesting thing about this news is that it's so easy to get caught up in this larger conversation of the battle between Google and OpenAI, and this larger phenomenon of competitive accelerationism, that we don't stop and remember how remarkable these new features are.

11:46When it comes to increasing the utility of ChatGPT in a day-to-day way, The ability to interact going back and forth via audio makes it unbelievably more useful for a mobile world. But the ability to use images as inputs, especially when on the go, makes ChatGPT so much closer to the actual super-powered AI assistant that so many people have imagined. The bike example may seem small, but that's the type of thing that people interact with every single day, day in and day out. That's the type of thing that people use Google for. I wonder what percentage of my Google searches have something to do with finding instructions or how to do something.

12:19It's probably a fairly big percentage, relatively speaking. By having this type of image input, ChatGPT is effectively competing not just with Google searches, but with my FaceTiming my brother who's much more technical than I am to have him try to figure out something for me. We talked a bunch last week about how Google is trying to differentiate by just loading up on actual utility and making their AIs more useful through integrations with other tools like Google Workspace. And then in the wake of that conversation, we saw Microsoft integrating AI everywhere through Windows 11 updates. and now OpenAI expanding the sort of day-to-day type of capabilities that will make ChatGPT much more powerful.

12:54Now, this would be interesting if it was the only ChatGPT and OpenAI story, but it was not. However, from this part of all confirmed announced things, we are now moving wildly into the realm of speculation, so a huge grain of salt warning for everything that comes next in this discussion. Over the weekend, there was a bunch of discussion on Reddit from two users who claimed that they had access to OpenAI's internal models and who were sharing some of the information that they had seen. I'll link to the specific posts, or at least Twitter screenshots of those posts, but here are some of the highlights that these users claimed.

13:26One of those users, Feltsteam, writes, So OpenAI obviously isn't just slowly developing one model at a time, but are of course working on multiple. The one that I know most about has an internal name of Iraqi, so it is kind of wild. So far as I know, it's an everything-to-everything model, meaning you can input on any combination of text, image, audio, and video. So what are some of the other details that these posters give? Well, one, they say that Iraqi succeeds GPT-4 capabilities and can match human experts in many different fields. They claim that hallucination rates are much lower than GPT-4.

13:56And interestingly, that half of the training data was synthetic. Now, this has been an ongoing conversation about the extent to which synthetic data might be problematic for training AI models in the future. Although there have been some results, including an unreleased Facebook Llama 2 model, that suggests that synthetic data actually can increase performance as well. In other words, that's a big open question, so it's fascinating that potentially this advanced model has 50 % of its data coming from synthetic sources. Now when it comes to when this stuff is coming out, the poster writes, In terms of release date, they originally didn't plan to release in 2024, but I think it's entirely possible to see it released sometime during 2024 as their timelines have been accelerated.

14:30Though it is their fault that everyone is accelerating in AI development as the release of ChatGPT and GPT-4 showed what was possible, and now people are slowly catching up, so it's complicated. Now, the other speculation around OpenAI comes around a Twitter account called Jimmy Apples. People are paying attention to this a little bit more than they might a random Twitter account, because after the information reported on September 18th that OpenAI was going to be releasing these multimodal features, and that they were working on a new multimodal LLM called Gobi, Jimmy Apples pointed out his own tweet from April 28th, where he said the big multimodal currently in the works at OpenAI is called Gobi.

15:03Should I leak more? Given that they were right about that, people are paying attention. And on September 18th, Apple's tweeted, AGI has been achieved internally. Now add to that a bunch of cryptic tweets from Sam Altman. One being, sure, 10x engineers are cool, but damn those 10 ,000x engineers and researchers, dot, dot, dot. And the other being, short timelines and slow takeoff will be a pretty good call, I think. But the way people define the start of the takeoff may make it seem otherwise. So this is just ratcheted AGI speculation to about a thousand. Simian CPS tweets, Can we consider seriously the hypothesis that 1.

15:36The recently hyped tweets from OA staff, 2. AGI has been achieved internally, 3. Sam Altman's comments on the qualification of slower fast takeoff hinging on the date you count from, 4. Sam Altman's comments on 10 ,000x researchers are actually mapping to something true? The implications are so crazy in terms of power shift or levels of risk over the next few months. Now, Sully Omar captured some of my feeling when he wrote, this whole thing is giving weird vibes. There's two possibilities. One, OpenAI has achieved AGI internally. Two, they're messing with everyone slash hyping things up for fun?

16:06But one thing is for sure, AGI is coming way faster than everyone thinks it is. Eliezer Yudkowsky seems to agree. The Twitter account at PauseAI responded to Sam Altman's tweet about short timelines and slow takeoff and said, how about we abort launch? To which Eliezer responded, you're talking to the wrong person. OpenAI has zero ability to stop the avalanche they started. That's now a matter for treaties between major powers. So friends, lots of intriguing things happening in the world of AI. We've certainly got another example of that competitive accelerationism we've been talking about. And holding aside all of the speculative stuff about AGI, we have a massively more performant and useful ChatGPT coming right around the corner.

16:44From the sheer standpoint of people who use ChatGPT and tools like MidJourney for productivity, October is gearing up to be a very, very good month. We will, of course, keep you up to date on all of the developments, including probably any relevant speculation here at the AI Breakdown. But for now, that is going to do it for the show. Thanks, as always, for listening or watching. And until next time, peace.

From the publisher

The race towards multimodal LLMs is heating up! With rumors of a big impending launch of Google Gemini, OpenAI is racing to push out their multimodal features. Today they launched the ability for ChatGPT to carry on audio conversations, as well as to use images as inputs. Before that on the Brief, Amazon to invest up to $4B in Anthropic.
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI. 

Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe

Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown

Join the community: bit.ly/aibreakdown

Learn more: http://breakdown.network/

More from The AI Daily Brief: Artificial Intelligence News and Analysis

All 1,099 episodes
ChatGPT Can Now See and HearThe AI Daily Brief: Artificial Intelligence News and Analysis · 17 min
Listen in VO