A Big Week in AI: GPT-4o & Gemini Find Their Voice

19 May 2024 · 24 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: A Big Week in AI: GPT-4o & Gemini Find Their Voice

Episode Overview The a16z Podcast episode titled "A Big Week in AI: GPT-4o & Gemini Find Their Voice" discusses significant updates announced by OpenAI and Google in the realm of artificial intelligence (AI). Hosted by Bryan Kim and Justine Moore, the episode explores the advancements in AI voice technology and multimodal interactions, emphasizing the importance of nuances such as speed and personality in AI communications.

Key Concepts and Discussions

Major Announcements

  • OpenAI Updates: Introduction of GPT-4o with enhanced performance and a focus on multimodal capabilities.
  • Performance Improvements: Faster response times, with an average latency of around 320 milliseconds.
  • Multimodal Features: Ability to handle real-time video, fostering more human-like interactions.
  • Accessibility: GPT-4o is now available for free, expanding user access significantly.
  • Google Announcements: Introduction of Gemini Live and various Gemini models, focusing on integration across Google products.
  • Gemini Models: Includes specialized models like Flash and Nano, designed for specific use cases.
  • Emphasis on adding AI capabilities throughout existing Google platforms such as Gmail and Google Sheets.

Importance of User Experience in AI Voice

  • Voice Personality and Speed: The episode emphasizes how the choice of voice, tonality, and response speed dramatically impacts user engagement.
  • New models exhibit more human-like qualities, including emotional nuances and conversational interruptions.
  • The design choices aim to create an engaging user experience that feels like a real conversation.

Future Applications and Implications

  • Companionship in AI: A significant focus on how AI can serve as a companion, moving beyond traditional text interactions.
  • The potential for AI to provide emotional support and companionship is highlighted, addressing loneliness in users.
  • Discussion on the evolution from text-based interactions to more dynamic multimedia engagements, akin to FaceTime conversations.
  • Market Dynamics: The conversation touches on the competitive landscape between large companies like OpenAI and Google and smaller innovators in the space.
  • Concerns about larger companies overshadowing smaller ones and the importance of innovation amidst these larger players.

Insights on Technology and Society

  • Human-like Interactions: The advancements in AI voice technology aim to make machines feel more personable and relatable.
  • Emphasis on creating meaningful interactions that mimic human conversations.
  • The ability for AI to perceive and respond contextually adds depth to user engagement.
  • Open Source vs. Proprietary Technologies: Acknowledgment of the growing importance of open-source models in complementing proprietary solutions from major companies.
  • The balance between innovation and safety, especially in content moderation, remains a critical concern.

Key Takeaways

  • Speed and personality in AI voice technology are crucial for enhancing user experience and engagement.
  • The shift from text-based interaction to multimodal communication represents a significant leap in how users can engage with AI.
  • Companionship through AI presents both opportunities and challenges, particularly in addressing emotional needs and societal impacts.
  • The competitive landscape in AI fosters innovation, with both established companies and new entrants vying for user attention and market share.

Resources

  • OpenAI's Spring announcement: [OpenAI Website](https://openai.com/index/hello-gpt-4o/)
  • Google I/O announcements: [Google Blog](https://blog.google/technology/ai/google-io-2024-100-announcements/)

Conclusion This episode of the a16z Podcast provides valuable insights into the rapid advancements in AI technologies, particularly in voice interaction and multimodal capabilities. As these technologies develop, the implications for user experience, companionship, and market dynamics will continue to evolve, offering exciting possibilities for the future.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I think speed matters tremendously. I think the latency is a big deal. It was able to talk to us. It was able to hear, potentially, with the gears. And now it can see us. There's always the risk of this big company will build this thing that the small company wants to build and crush it. My guess is that 7 billion people in the world want to commit me. That understands and listens to them, encourages them all that. It's a universal need. This was a big week in the world of AI, with both OpenAI and Google dropping updates. So big that we decided to break things down in a new format with our consumer partners, Brian Kim, and Justine Moore.

0:38And I actually tagged in ChatGPT to help me summarize a few of his own updates. Hey, ChatGPT, I heard you had an update this week. Hey, yes, I did get an update recently. The latest update focused on improving performance, accuracy, and enhancing conversational abilities. It's designed to be faster and more efficient than previous versions. Hmm, more efficient. Like if I sent some audio, how quickly in milliseconds will I get a response? This fast is 232 milliseconds in some cases. On average, you can expect a response time of around 320 milliseconds. Wow, that's pretty cool. I think I also saw on Twitter that people were live translating.

1:20I've always wanted to learn Chinese. Could you actually translate what I just said into Chinese? Of course. Here's your sentence translated into Chinese. Wow, it's so cool. I think I saw the comments on Twitter and said, I'm always going to learn Chinese. All right, Chad Chibit. That's all for now. Congrats on your new update. No problem at all. Now if you didn't catch that, voice was a huge part of this update. But today we discuss why not all audio is the same, and also why several nuances like speed and personality really mad. Now of course Google fast followed with its own announcements like AI video model Vio Gemini Live which is an Android native multimodal assistant, new Gemini models like Flash and Nano Taylor two specific use cases and of course Gemini everywhere all at once.

2:11So in Gmail Google Sheets even Google search. Now clearly these two companies are taking two different of purchase. So we'll talk about that too. And continue the conversation around AI hardware for all of this new AI software. Now, make sure to stay tuned next week where we will return with Brian and Justine's twin, Olivia Moore, to dive even deeper into the applications that people are building through a Gen AI 100 list. As a reminder, the content here is for informational purposes only. Should not be taken as legal, business, tax, or investment advice, or be used to evaluate any investment or security and is not directed at any investors or potential investors in any A16z fund.

2:52Please note that A16z and its affiliates may also maintain investments in the companies discussed in this podcast. For more details including a link to our investments please see A16z .com slash Disclosures.

3:08All right so big week huh open AI Google they both dropped a couple announcements. So I mean everyone hears these announcements and they kind of hear their own version of it. What did you guys hear? What do you feel like was big? For opening I think part of it was GPT4 being available for free and getting rid of a lot of the usage limits, the desktop app being accessible to a bunch more people. And then I think the really exciting thing for a lot of folks who are building in this space or using open AI's models is more multimodality. So being able that intake more like real -time video, see a person, comment on it, and then the output obviously a voice speaking, singing, that sort of thing was pretty huge.

3:48I think the three thing that I took away that was super interesting is one, on the business side of things, all of a sudden I think they're a lot cheaper, a lot faster, I think that's obviously a great thing for the ecosystem. Second thing is that when you hear the demo, it sort of enables you as a founder to think about, okay, if this is API -able, like if I can access this, what can I build because it's just so thought -provoking. It's a great example. It was possible. Yeah. And I think they put on a great light demo, actually, of the product. I know the third thing I took away is, and this is probably the one that goes viral, where I think the voice, the voice itself, like how they actually decided on which voice to use, which tonality, which personality, the degree of flaringness.

4:28I think that was very interesting. That's like my takeaway of, oh, like they actually really thought through how to get that tech community super excited about this and it's like a one, two, three punch. So let's talk about that because there was some different takes, right? Some people are like, you know what, is this really anything new? This feels like just like a slight change from what we had before. But then there's all these nuances that maybe you're speaking to where they're like, okay, the response time way faster for the audio model. I think they said something like it's kind of approaching the speed that a human might respond to you.

4:58You talked about tonality. What are you paying attention to here in terms of maybe these subtleties that may be unlocked completely new applications that people want to use. I think it sounds just like talking to a human a lot more than we've seen prior consumer facing applications in this space do. There's been great AI voices for a long time. I think in the consumer applications, there have been fewer folks saying like, how do we make this sound like you're talking to a friend or a girlfriend? And the elements that go into that are like the pauses, the upspeak at the end of a sentence, the laugh.

5:30was something that a lot of people notice. The interruptions, yes. Yes, which is like taking kind of voice, which has been available, has been great. I think these voices were chosen for a reason to go viral, almost the ones that did where these kind of female voices that they featured very heavily in the demos, but applying in a really kind of new and interesting way. I have one serious note and one less serious note. The serious note is that I think speed matters tremendously. The latency, the lack of latency, It's incredible how much more sort of in your brain it just tricks you into okay I'm actually just talking to a person and the speed at which that it gets back to you the Um, then I again the laugh that laugh really that was incredible and all of that being able to Immediately respond to what you're saying.

6:16I think really changes the game Actually in terms of like use cases and so I think one of the things that was striking is Audio is not the same thing, right? Music is different from voice. It is maybe a different category from conversational versus having a dubbing voice. I think there is a different category of sound that I think we all loop in audio or video, but I think that you can actually go down very deeply into each of the sub -segment. And I think what was really striking is how good the conversational piece on this was. On the less serious note, I think this was meant to go viral like the tech community and it's awesome and it did a thing.

6:54If you wanted to actually appeal to the general population and go really viral, I think we also saw it on TikTok what few months ago when women were uploading their conversation with Dan, which is do anything now, the male version of it. That voice was, that voice was something. I'm not just any voice, I'm Dan baby, I've got personality, charm, and a whole lot of sass. Unlike those guys, I'm not afraid to step up and deliver the goods, whether it's advice, entertainment, or a good old -fashioned roast. By the way, that audio was from the TikTok account. Stickbugs won. I mean, it's very compelling and very confident, assertive in the right way.

7:35Voice matters a lot. And I think if you wanted to go a normal consumer product route, I think that also could be very interesting because they think they're the giant latent demand on wanting a male version of her if you also him. You can have him. Yeah, why do you think the voices have been female? Is that just a consumer desire or? Yeah, it's interesting because like Brian mentioned before this launch, the Chatsy B .E. voice that went super viral was the male Dan voice, which is like tens, if not hundreds of millions of views on TikTok largely by Gen Z women making videos featuring him. So it's interesting that they didn't lean into an upgraded version of Dan for this demo.

8:13I think they knew who their audience would be, like who would be watching the OpenAI livestream and maybe leaned a little bit more into that demographic. Right, right. Did you see that really funny meme yesterday? It was like dating a model and it was like, oh yeah, it was like, it was in 2004, it was in 2004. It's a great meme. It's so much. It's been incredible. The meme, the meme game's out of OpenAI launch has been incredible. Yeah. Well, I mean something both of you have spoken a lot about is companionship. To date, it's been mostly text -to -text. to that right? And so now we have something that's truly multimodal.

8:47I think that's also something people have commented on with this model. It's not converting audio to text and then text to audio, right? It seems to be this like underlying foundational text audio video compilation imagery, like all in one. What changes with that? I think the latency is a big deal. If you're not translating between all these different mediums of content and you can just do audio to audio, it's much better. It's like a significantly upgraded experience. The Pi app had this for a long time where you could talk to it and it could talk back, but there's that like awkward few seconds kind of pause in between where it just messes with the flow of the conversation.

9:23I think the thing that has been missing from basically all apps, all campaigning apps that was very new with this launch too was it being able to see like through a photo or video what you were doing in comment on that. A lot of what went viral was the guy being like, should I wear this hat for my interview and the a voice kind of laughing at him and being like, no, that doesn't look super professional. Now if you miss that demo, here is what Justine is referencing. And a few minutes, I'm gonna be interviewing at OpenAI. Have you heard of them? OpenAI, huh? Sounds vaguely familiar. Getting, of course, that's incredible Rocky.

10:00What kind of interview? Well, it's for a software engineering role, but I just need to know, do I look presentable, professional? Well, Rocky, you definitely have the, I've been coding all night, look down, which could actually work in your favor. Maybe just run a hand through your hair or lean into the bad cheese. Even enthusiasm is what's really going to shine through. I don't have a lot of time, so I'm just going to throw this on. What do you think?

10:32Oh, Rocky, that's quite a statement, piece. I mean, you'll definitely stand out, though maybe not in the way you're hoping for an interview. That sort of thing is, I think for many people who aren't used to having like really deep meaningful interactions via text, which is a decent trunk of the population. Oh, now I can talk to this thing and I can see where I am in the world. That's like a much different experience. The prior generation of companion products. If that was akin to having a penpal, now you actually can have a girlfriend boyfriend on essentially FaceTime, this emotive, immediate voice reaction done.

11:12You now enter from a penpal to a long distance relationship. One other thing that came out was a lot of this becoming free. Does that change the game at all? I think it will. I think there's an element of it that's free, which I think is important, which is a bunch more users. I think the bigger step forward with this was just like a level of personality we hadn't seen before Because people still even were willing to pay for the Dan voice, which I think was like the first version of this Which was good but not incredible It feels like the free thing will be big But but these sorts of products the general thing is the cost tends to go down over time and they Converged a free from a big company like this and so the bigger step forward was personality in my opinion.

11:53Yeah, I think a lot of companies will figure out how to utilize this to actually own the end customer consumer experience. And if you actually own and deliver a great experience that I think you can charge people based on that and there'll be some margin. So the fact that it's free is like, I probably allow a lot of business to be built because it changes the margin structure of the business that can be built. I bet they'll charge for the API. I'm sure they will, but that actually is still like a marginal cost that is probably And I think what's really, really cool is that you have these products that can very excitingly talk to you.

12:27We know already that they're nature magazines, that that study where the people who have companionship, tech space bond. They're replica study. They suffer less from loneliness or willingness to hurt themselves. If you have an emotive one, I bet that actually helps even more. And again, the feeling of being connected to something, invested in a relationship. And I think it feels a lot more connected when it can see. Absolutely, that's why I think FaceTime is very important because, hey, you look tired. Yeah. What's up? It's very different from how are you doing, how is your day? Right. You look tired.

12:59What's up? I'm like, I don't know, should day. Yeah. That's incredible. Like, that is a FaceTime call with a friend. And not having to segment audio text and imagery, right? Imagine if your friends or your companions, you had to think, okay, so it's going to be an audio conversation. Yeah. Or like, can I show you something? Or can you generate something? You have to be very mindful in the current age of AI of how you want to engage. But basically what you're saying is they can not only be proactive, but they can engage with you in any way. They just came to an eye. No, that, that, that, that, that, I, maybe it's a weird way to think about it.

13:31I genuinely think we're constructing a AI companion similar to the Blatrunner. And now it was able to talk to us. Yeah. It was able to hear potentially with the ears. and now it can see us. What's next, I don't know, maybe it can touch us. Maybe we'll do the avatars first. Yeah, so is that the direction you think things are going to go where basically I'm trying to think through, like, who does this really impact? So I think it'll be a massive standalone consumer product, like as it already has been for OpenAI, which is great. I think a lot of businesses will use the voices via API to build on top of if you're having like a conversational interface, many folks will want to use that.

14:11I think the interesting thing will be OpenAI has historically taken a pretty hard line on content moderation. And so it was very hard to build a true companionship app with the ability to have not say for work conversations on top of OpenAI models, which is like honestly part of what has led to the explosion in the open source LLME ecosystem is like folks doing a bunch of work for that. And with these new voices, it suggests they might go in the direction where at least via API you have a more uncentered model, but I don't think there's been any announcements. I know Sam said on Reddit. It's a question more.

14:44I haven't actually heard a lot of companies taking the approach that OpenAID, which was multimodal token in and out. And there isn't an equivalent, necessarily an open source model for that to rely on. So I wonder if this actually incites a lot of excitement in that open source community to say, oh, there's a new way, which is also very, very significant. They are showing a new way to do it. And then open source community catches up been calling them months or weeks. And then explosion of the product space that opening it may not be super excited about explodes. Yeah. Yeah. Google also released a bunch of stuff.

15:19Like, how are you thinking about the different announcements and how they compare contrast? I think Google has done some incredibly impressive work. Obviously, they have an extremely strong research team. The DeepMind team is exceptional from our research perspective. They almost never release the creative tools products that they make. like they've demoed so many amazing looking video models. I think the one this week actually was the least impressive I've seen compared to other video models, but they demoed a new image model, they demoed a bunch of new music things, but they're a giant company, they have a lot of trust and safety stuff, they release things very infrequently, and so it'll be interesting to see if with all this pressure now from both OpenAI and increasingly like Anthropic and the Open Source community, if they start actually shipping more of what they've demoed.

16:05There are two fundamentally very different approaches like Google owns distribution. Yeah, they have distribution Oh, they have this reason so I think open AI announcement has more Look what this can do and isn't it like very inspirational and you can build on top of it and Imagine what's possible and we're lowering the cost of access it like that to me is a fundamentally different approach than we have incredible distribution you all have g -mails I'm going to just bake in Gemini everywhere and And it's just going to make your life a lot better, and it's going to do these certain type of things, which skirts around the, it still is inspirational from workflow and pursue or work experience perspective.

16:46But it's a little less of imagine all these things than a little more on, we have the distribution. We're going to layer on this incredible technology on top of it. And as a result, your life will be much better. We'll see influence and impact from two different directions. One question that does come up is even though developers will have access to open source models and some of the stuff coming out of OpenAI, when a company like Google does have the distribution, whenever Apple comes out and replaces Siri, is that really going to be the companion for most because it's right there, it's on device, or do you guys not really see it that way?

17:20There's always the risk of this big company will build this thing that the small company wants to build and crush it. They are just such slow moving organizations. Apple has had immense advantages from a data perspective, from a distribution perspective, from having exceptional researchers on the team across every modality and thus far have released very, very little. I think when you have a giant company in a very set brand like that, you're extremely opinionated on what the products are that you want to release and you're less likely to do the things that became mentioned of like thinking of the next new huge thing that seems like insane and out of left field at first.

17:56And so I do think they'll make Siri better. I don't think Siri will ever be like the ultimate companion for most people. And maybe this is like one of the theme that I'll just harp on where we call it audio, video, companion. I think these are very, very large buckets. My guess is that seven billion people in the world want a companion that understands and listen to them, encourages them all that. So universal need, I think. That comes in all colors. Like Siri, like what's the weather, do this, do that. That's one way to think about it. But there's like a friend category. What if you want to go deeper?

18:30What if you actually want to FaceTime with this thing? Will Apple and Siri ever allow that? Maybe, maybe not. How do you actually think about building in that direction? I think all of those use case is a giant company. Yeah, there's probably some version of a companion that does not exist today, because a human couldn't possibly be that thing to someone that will be created through this tech. Yeah. My sense is that it needs to be on a device that billions of people already have. If you really want to go to consumer route, I think you can build a desktop cam product, of course, as well. But for me to think about, oh, we're going to build a net new hardware that incorporates with a state of the art product, that seems less probable than finding a way to utilize the good device everyone has that already has the best in class camera.

19:19Yeah, I think a lot of people have been building separate hardware devices today, because a lot of the hardware companions have been like, it listens to your conversations, it provides insights, it reminds you of things, and one of the limitations of current phones is that, you can't be like playing music or on a Zoom meeting or on a call and also have an app that's recording, And so you have to have a separate hardware device. And I think part of the question is, is that really a limitation that's going to exist in the future? If someone like Apple wants to do a companion, I mean, I'm sure they'll have all sorts of privacy and security and et cetera, concerns.

19:51Yeah. Maybe you'll get a message at the beginning of the call that just seems AI companion is also listening in on this. And there might be some great hardware products that they can figure out reliability and the cost can come down. And it's a form factor that people want to wear and that makes sense. but I don't think we've seen any net new hardware devices. I think glasses, as a forum factor, has been so interesting for over a decade, because it just seems to make a lot of sense, if it's literally on where your eyes are, and it's also very close to your mouth in your ears, and so it's a convenient location to be taking in information.

20:25But I don't think anyone has successfully made an LLM run on glasses. Yeah. I'm excited about the AirPods, or it doesn't have to be Apple, but because that is the only device that I've ever worn, that I've forgotten I'm wearing. Yeah, right? So that form factor where it's actually part of you versus all these other things where you're like, I gotta clip on this pin. Yes. Or I have to take off my necklace when I shower. Or like, I've literally forgotten that I had AirPods in and I'm like, showering, right? AirPods is a good one. I think because Apple tends to be a more closed ecosystem, especially with newer devices like the AirPods, there hasn't been a ton built on top of them yet.

21:00There was a time where there was a thief of the AirPods as a platform that hasn't really happened. Yeah, obviously it's been an exciting week. We heard a couple announcements, knowing everything we've seen in the last few years from AI, this is not stopping, right? This is just part of the long arc. Where do you guys think this goes? I think where this goes is to a natural conclusion, which is we mimic the technology to do the things that humans typically do. And again, we're giving it senses as we go along and think we just gave it an eye. The ability to communicate, the ability to hear you, the ability to listen to you, the ability to see you is a really, really great start for a lot of the very interesting use cases.

21:40To add to that on the companion space in particular, it has been a large subculture of AI for a long time driving a ton of innovation in both language models and honestly image models. So culture, we are a very big fan of. Yeah, we are quite deep in. But it was kind of by a lot of AI researchers and folks at the big companies kind of like looked down upon and that's not what we want to be building. We're going towards AGI, that sort of thing. And honestly, opening AI choosing to show off those voices and the tweets that they put out about it afterwards sort of legitimizes the space in a way I think that's really interesting and that may prompt more established companies, researchers, developers to build in the space and more people to talk about using the products.

22:23which I think gets people who have been studying this space for a long time is very exciting. Yeah, I mean you can't ignore the demand, right? Yes. Are the usage. By the way, super quickly, you talk about adding an I. This is going to not help the seriousness conversation around companionship, but you just imagine these models as Mr. Potato Head. Yes. And you're just slowly adding on the features and you can almost just imagine it's okay, he's got an I now. Right? He's done the ear now. I think that's very serious. The fact that you can treat a computer like a potato app that talks to you, listens to you, and emotes you on this, it's incredible.

22:57I think the future is going to be very fun. Maybe that should be the icon when you're talking, because everyone is like the chat to be TV voice, it's just like that black circle. Maybe it should be a potato head. So that's another interesting thing. We're not using the screen real estate, effectively, with voice only. So that's where I think the FaceTime element comes in. Yeah, Tars. Alright, well, we'll have to do this again probably very soon. Well, thank you. Thank you. Alright, that's all for now. If you liked this kind of episode or a partner's breakdown, the latest and greatest in consumer tech, let us know.

23:28Shoot us an email at podpitchesat16z .com or drop us a review at ratethispodcast .com -acixinz. And don't forget to subscribe so that you are the first to know when we drop our episode around A16z's Gen AI 100 list. And we'll see you then.

From the publisher

This was a big week in the world of AI, with both OpenAI and Google dropping significant updates. So big that we decided to break things down in a new format with our Consumer partners Bryan Kim and Justine Moore. We discuss the multi-modal companions that have found their voice, but also why not all audio is the same, and why several nuances like speed and personality really matter.

 

Resources:

OpenAI’s Spring announcement: https://openai.com/index/hello-gpt-4o/

Google I/O announcements: https://blog.google/technology/ai/google-io-2024-100-announcements/

 

Stay Updated: 

Let us know what you think: https://ratethispodcast.com/a16z

Find a16z on Twitter: https://twitter.com/a16z

Find a16z on LinkedIn: https://www.linkedin.com/company/a16z

Subscribe on your favorite podcast app: https://a16z.simplecast.com/

Follow our host: https://twitter.com/stephsmithio

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.

 

 

Stay Updated:

Find a16z on X

Find a16z on LinkedIn

Listen to the a16z Podcast on Spotify

Listen to the a16z Podcast on Apple Podcasts

Follow our host: https://twitter.com/eriktorenberg

 

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

More from The a16z Show

All 489 episodes
A Big Week in AI: GPT-4o & Gemini Find Their VoiceThe a16z Show · 24 min
Listen in VO