Gustav Söderström: Catapulting Spotify to the Front of the AI Revolution (Encore)

12 Sep 2024 · 1 h

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Generative Now - Episode with Gustav Söderström

Episode Overview Title: Catapulting Spotify to the Front of the AI Revolution (Encore) Host: Michael Mignano (Lightspeed Partner)

Guest

Gustav Söderström (Co-President, CPO & CTO at Spotify) Description: The episode revisits a conversation exploring Spotify's evolution as a consumer AI company, its recommendation systems, and future directions in AI and media.

Episode Chapters

  • 00:00 - Introduction to Gustav Söderström
  • 01:29 - Spotify as a pioneer in consumer AI
  • 05:57 - Philosophical shift in UI design for AI
  • 12:52 - The importance of dialogue in development
  • 15:00 - Building for future technology rather than existing tech
  • 19:43 - AI's impact on media creation
  • 25:00 - Challenges in model training
  • 29:45 - Understanding AI through anthropomorphism
  • 37:04 - The consequences of automating cognitive labor
  • 40:42 - Potential intelligence growth of AI
  • 47:45 - The future of AGI (Artificial General Intelligence)
  • 53:48 - Addressing concerns about AGI

Key Themes and Discussions

Spotify's Journey to AI

  • Initial Focus: Spotify began primarily as a music curation service, leveraging user-generated playlists for data.
  • Transition to AI: The company transitioned to a recommendation-focused approach around 2015, utilizing machine learning to enhance personalized experiences.
  • Significance of Data: Spotify's unique dataset from user-created playlists is a powerful resource for understanding music preferences.

UI and AI Integration

  • Philosophical Shift: The user interface (UI) is evolving to support AI, rather than the traditional model where AI supports the UI.
  • User Interface as an AI Product: There's a growing emphasis on creating interfaces that provide clear signals for AI learning, shifting the focus from dense UIs to formats that enhance AI’s understanding of user preferences.

The Future of AI and Content Creation

  • AI in User-Facing Products: Söderström discusses the expansive role of AI in improving user experiences, including podcast and content recommendations.
  • Predictions for Media Creation: Anticipates a future where AI enhances the productivity of content creators on platforms like Spotify, rather than replacing them.

Action Models and AGI Concerns

  • Emerging Action Models: Discussion about the potential for AI to take actionable steps in a user’s environment (e.g., automatic recommendations based on context).
  • AGI Definition and Risks: The conversation touches on the philosophical implications of AGI and the balance between AI intelligence and the risks it may pose to society. Söderström expresses optimism that AI will evolve to understand human needs and contexts better.

AI’s Impact on Employment

  • Cognitive Labor Automation: There is concern about how automating cognitive tasks will affect jobs but also a recognition that AI could improve productivity rather than reduce the workforce.
  • Future Workforce Dynamics: Companies are more likely to invest in enhancing existing talent with AI tools rather than replacing them.

Ethical Considerations and Responsibilities

  • Training Data Challenges: The "wild west" of data usage for training models is highlighted, with a call for better regulatory frameworks.
  • Empathy in AI: Söderström discusses the possibility of AI developing empathetic responses based on its training on human interactions, suggesting a more caring approach to AI development.

Conclusion Gustav Söderström presents a compelling vision for the future of AI at Spotify, emphasizing the importance of thoughtful integration of AI in user experiences. His optimistic outlook encourages a balanced view of the technological advancements and their implications for society, recognizing both the potential benefits and the necessary precautions.

Call to Action Listeners are encouraged to engage with the podcast and provide feedback, following Lightspeed's updates on various social media platforms.

--- For further information, follow Generative Now and Lightspeed on their respective platforms.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:04Hey, everyone, and welcome to Generative Now. I am Michael Magnano. I am a partner at Lightspeed And for this week's episode, we are revisiting one of our favorite conversations with a big name in the world of consumer AI. It's Gustav Soderstrom. He's the co-president, CPO, and CTO at Spotify. And he's one of the leading minds in all of AI. And it definitely showed during this conversation. We covered Spotify's evolution towards consumer AI products and how close we are to action-based agents and his take on AGI. So take a listen to this conversation with Gustav Soderstrom. Hey, Gustav. Hey, Mike.

0:43Great to see you. Good to see you. So normally we'd catch up, say hello, you know, small talk. But you and I, when we get talking, we could talk for hours. So we're going to dive right into it so we don't waste any time. So I want to talk about Spotify a little bit before we get into some more general AI stuff. But, you know, people don't really think of Spotify as an AI company. But I kind of think of it as the first true AI consumer product company. Give us a brief story of the history of sort of machine learning driven personalization at Spotify. I like it. I like it. I think there's some truth to this in the sense that Spotify started like most services in sort of late 90s to early 2000s or almost 2010s as a curation service.

1:33So the name of the game back then was to take some good, like books or music or movies, digitize them, and then get users to sort of catalog them for you or organize them or curate them, if you want to use a fancy word. So that was Facebook, you digitized your friends, and then you got people to curate the friends into groups and graphs. It was Amazon with books and so forth. And Spotify was the same. The catalog was digitized by the MP3, and then the MP3 was sort of separated from the CD. There was Pirates, obviously, and so forth. And so music was digital, but what Spotify did, it asked users to curate these things into playlists, right?

2:11And so the user was doing work for themselves to create a great music session, but they also helped other users by creating these playlists because they were public and shareable, right? But what also happened was as the users curated these playlists into groups, it also generated a lot of data about what songs go together. And pretty early on, I joined Spotify in late 2008, 2009. There were already some people at Spotify who knew machine learning. Back then, it wasn't so much neural networks. It was more collaborative filtering and so forth. But already back then, some people were experimenting with taking these playlists and see if you could create, using collaborative filtering, vector spaces that could describe these similarities between tracks.

2:54So really the training, the playlisting was the training data for Spotify. And, you know, already tens of years ago, we had billions of these curations of tracks that go well together. So in that sense, we happen to have a lot of training data without really that being the goal. But to the credit of these people, they realized pretty early on that this was great training data. And so we started investing in recommendations. recommendations and at first these recommendations were a support to the curation right so the first thing you would see was would be similar artists on the artist page right so it's kind of a hidden feature but we could see that people really love these similar artists things they would just click through the graph of similar artists and so then more and more we started saying like well if we understand which artists you might like maybe we could just suggest some tracks from these artists that you might like as well and we sort of pivoted from from curation into recommendation And that was fortuitous for us because the technology kind of developed at the same time.

3:55But in that sense, we were actually pretty early before the whole sort of maybe 2015 machine learning wave that happened. And we started sort of pivoting the company from curation first to recommendation first. Now, obviously, you can still curate. You can still create your own playlist and so forth. But more and more, the core promise of Spotify is the personalization and the recommendations. And if you think about it, that makes sense for us because music in itself is a commodity product. You can get the exact same tracks on any music service for roughly the same price. So it's very important to us that we bring some additional value.

4:32And more and more, that value transition from being curated playlists into being recommendations and how well we understand you. That data must be so, so powerful and so unique. If you think about some of these other services that are, you know, recommending content algorithmically, they're doing it by making a guess based on what people are maybe watching for more than they're watching something else, like how long they're looking at the screen. But in the case of Spotify, these users specifically said, I want to listen to this track next to this track. It feels like one of the most powerful data sets there could be for media.

5:07Yeah, I think that's right. I mean, in most other spaces, movies and even podcasts for us, you have the consumption signal, like someone listen to this and then listen to that. But that's less powerful than this explicit curation. Someone saying literally like these five tracks, they go together, you know, because it's something. And then we can try to understand what that something is. So it's a very explicit data set. Yeah, it's really cool. So AI has in many ways then been sort of in the background of Spotify for a while now with the curation, as we just talked about. But it's sort of making its way into the foreground now with things like AI DJ, some of the new playlists that you've come out with are even more leaning into AI.

5:52What are some of the other areas of the product that you're going to be applying AI to sort of on the user facing side? So, I mean, I think you can look at it practically and look in all the places where we're going to use it. And you can also think about it philosophically. I think if we start actually philosophically, I think the way to think about it is exactly what you said. When we started with AI, as I said, it was in support of the UI. Like the AI was supposed to help the UI, but the UI was the product. At least that's how we thought about it. And recently this is switching. Now the UI is there to help the AI and the AI is actually the product.

6:30And you think more and more about it. What is it that we're trying to do? Well, we're trying to build some sort of approximation of you, really, and your tastes in music, podcasts, and audiobooks. And we're trying to understand and predict how you would react to this track, this podcast, this book. So we're trying to model you. And that model is actually inside the weights of the neural networks. And that is the real product. That is what you're paying for. So we started saying internally that the AI is the product. And the UI is there to help the AI as much as possible and help the user, obviously.

7:03So you want the UI to help give you a strong signal as possible from what the user actually wants to hear or watch or listen to or read. And so that's a shift, like a mental shift. And I've said this before, I think the product that sort of pioneered this was probably TikTok, where you can literally see that the UI is really just a full screen video feed. It's designed for the AI. It was 100 % in service of the AI. So that's not a shift that we came up with. It's a shift the entire industry is going through, but it's good to be explicit about it. So that's at a very high level. And then you can geek out about what it means at the end of the road as we're trying to model basically a human and its tastes and instincts and so forth.

7:50But then if you look at it practically, what that means is you want user interfaces that are better at giving you clear signal. So if you think through two types of interfaces, one is sort of the Netflix and former Spotify user interface where you have lots of little cover arts all over the screen and you scroll through them. In a UI-first world, that's pretty useful for the user because then they can evaluate lots of cover arts at the same time. So it's efficient. And you talk often about dense UIs. Lots of information on the screen because from a user perspective, it means the user has to scroll less.

8:27You can see many things at the same time. But then if you try to think of it from the AI's perspective, you know, there is this blog post from Eugene Wei called Seeing Like an Algorithm from like 2022, which I really like. So if you try to think you are the AI, that UI is very difficult because you can't see what the user actually understood. Did they look at all the items? And if they look at them, did they understand what that little square cover actually meant to represent it? Did they evaluate it and say, I'm not interested? or did they not even see it and you should actually show it again?

9:00So if you go to the TikTok UI, it's less efficient for the user. You only see one thing at a time, and to go through 10 items, you have to do 10 swipes. But from the AI's perspective, it is 100 % certain that you saw it and if you liked it or not. If you didn't like it, you would swipe. If you liked it, you would stay on it. So it's actually less efficient for the user potentially in the near term, but it's more efficient for the AI, which makes it way more efficient for the user the next time because it understood the user better. So that's a general way of like, how do you rethink UIs in an AI world?

9:32We try to make sure that the algorithm understands what the user preferences are more than if you didn't have an AI algorithm. So that's one place. And then we also obviously use AI all over the place. The content understanding is a big one. We used to have a problem of recommending podcasts before they got popularity because we didn't know who to recommend them to. or if they were safe or unsafe. Now you can use LLMs, you can machine listen to them and classify them both in which category they are, which, you know, create a vector for them to recommend them and also understand if they are safe for recommendations in advertising or not.

10:12So content understanding is a big, big piece of that too. And then obviously the recommendation algorithms themselves are getting much, much better with these very large embeddings from LLMs versus the much smaller ones that you used to do in the age of collaborative filtering. So that's some of the areas. Just to go back to the UI point really quickly, it strikes me that Spotify would have a little bit of a harder time doing this than other products that have visual media because so much of the consumption and the experience of Spotify is happening in the background. The user has the phone in the pocket or they're listening to Spotify in their car.

10:48What are some of the signals you can get when the user isn't actually interfacing with a visual UI? and how do you filter that back into the AI? Yeah, so that's a great point. And that's one of the challenges we actually have there. So if you think about someone listening to playlists in the background and maybe they're driving or something, right? So they're very, interactivity is very expensive for them. You have to bring up your phone, you have to unlock it with your Face ID, stuff that you shouldn't even do while you're driving, right? So then if you play a track and the user listens through the whole track, how do you interpret that?

11:22It could have been a track that they, absolutely loved. And it could have been a track that they hated, but not quite enough to skip through it, right? And it's very hard for the algorithm to understand. So it's exactly the problem of seeing like an algorithm. Now, fortunately for us in music, we have other signals like saves and playlisting, which to your point previously, are super explicit, not just that you liked it, but also that you liked it in the context of these other songs. So that's why playlisting has been so important for us to bootstrap this, because background listening is a weak signal to learn from.

11:53But this is also why you saw us investing in more foreground discovery mechanisms, literally sort of feeds where you can swipe through, you know, canvases and hopefully soon music videos of tracks one at a time. So, you know, as you are in, when you are in the foreground, we want more efficient formats to capture as much signal as possible when you're in the foreground, because in the background, it's harder to capture. So we're trying to make sure that we are efficient in the foreground, that we have the playlisting tools for really explicit curation. And obviously, we're also using the skips and so forth in the background, but it's a harder signal to learn fast from.

12:29This past year has obviously been crazy. We've seen so much innovation, so many new companies, products, technologies, and it seems like it's moving so fast. Every single week, there's a new model that's out benchmarking the previous model. If I think about that in the context of Spotify and, you know, my time spent there working with you. Obviously, Spotify is a super, super thoughtful company. You've always said, you've said it internally, and I've heard you say it publicly, talk is cheap, so you should do a lot of it. And what you mean, I think, is that, you know, you should really, really take your time and be really thoughtful and strategic.

13:04And I think that's reflected in a lot of some of Spotify's biggest initiatives. They were years, years in the making. How have you how have you managed that in this past year and maybe moving forward where things are happening so so quickly like how do you just keep pace in this new ai first world yeah so what i mean with this uh talk is cheap so you should do more of it is obviously it's always better to have provocative taglines because people remember them yeah uh i remember and and uh what i mean is it's uh it is expensive to to just talk and talk and talk if if you don't have like qualified people and and structured discussions but it's also true that it's much cheaper to talk than to build the wrong thing that that's what i want to get at and you can get surprisingly far by debating and discussing famously you know um some of the of the old uh greeks came all the way to the atom just from like deduction and reasoning that there should be something like an atom out there.

14:09So you can get really far by just discussing. That's my point. And I think there's been this other culture of move fast, break things, code decides arguments that is also obviously valid. There are points to that. But if you go too hard in that direction, you can waste a lot of time going really fast, absolutely nowhere. So you need both. That's why I want to push a little bit for like sort of Socratic debate and discussions with strong people. And this is the answer to your question. So one of the challenges now that things move this fast is that it used to be the case that some technology sort of came online.

14:45You could see it. You could test it. You could start saying like, oh, what products can we build from this? And he would build it for a year or something, and then you shipped it. And that was okay. now the tricky thing is like you have to predict where the technology is going to be and start building the things around it hoping that it matures if you want to be early you have to intercept the technology rather than wait for it and that makes it tricky so if you if you need to intercept it you need to have reasoning around what products are going to be able to exist you know six to twelve months from now um that doesn't exist yet and what do we need to build to be able to ship something then.

15:23Because if you start once it's live, you're going to be years behind the others. So one of these examples is, for example, the AIDJ, where if you followed LLMs closely, you know, the last like, basically all the way since the Transformer paper came out, you could see this trajectory. And we reasoned internally, you know, we listened to voice models and we said like, they're not very good yet. back then. But if this pace continues, they're going to be perfect somewhere between like 12 to 18 months, right? And so what would that mean? What products could you build if you had perfect voice? And on top of that, you could see the LLMs, which back in like GPT-2 were starting to be able to create good sentences, right?

16:12But they weren't that good. But then with 3, 3.5 and 4, you could do the same thing you could say like within something like 12 to 18 months we're gonna have things that can produce highly relevant personalized text and we're gonna have things that can articulate that text perfectly as if it was a human what product could we build and then we started thinking about this dream we always had like you know what if we could have hired like one dj per user you know someone who just knew mike really well and sat there all night like going through the catalog and set up a playlist for you. And then in the morning, it's like, hey, good morning, Mike, I've got this thing for you.

16:50And that would become possible somewhere within one to two years from where we were. So we started investing. And there was a lot of investment in infrastructure we needed to do to be able to do that, even though the models weren't there. And so we predicted that, we built a bunch of infrastructure, we started prototyping these products the text generation was not that good the audio sounded like a robot but it took us like 12 months and during that time exactly that actually happened and that's why we managed to launch the ai dj you know early last year which looked like you know right when it happened uh which you know i was very happy about but it was actually a year and a half in the making and it happened to coincide with when the technology matured.

17:38So it kind of looked like we saw it and managed to build it overnight, but it took about a year and a half. So that's what I think is necessary now, more reasoning and prediction about the future. So like you said, you had to sort of predict what was going to happen and you probably never could have predicted that it would have launched at the exact same time as ChatGPT or similar, but it looked genius in retrospect. I wonder what that experience now and maybe this whole AI wave that's happened over the past year has done to the culture of Spotify and how everyone thinks about building products?

18:07I mean, are more and more people starting to think this way? What does it do to the way people work? I think so. We're having more and more sort of what-if discussions and more and more discussions about not what can be done right now, but what we think can be done in six months or 12 months. And in a way, I actually think that, as you said initially, the pace of innovation is absolutely mind-boggling, right? It's just exploding. And that's true. But in a sense, I feel like the world has gotten more predictable, the AI world. Because during the last five years, when this happened, like the first few years, you didn't understand any of this.

18:50It was very unpredictable. Now, it is more powerful than ever. And it's certainly growing as fast as ever. But it's actually very predictable that it's going to get better. You have the chinchilla papers and so forth. So, you know, OpenAI managed to actually predict how good GPT-4 would be before they built it. And I think they know how good GPT-5 is going to be as well. And so in a way, the world is moving faster than ever, which makes it harder, but it's also more predictable now how fast it's going to move. Let's talk a little bit about media and content in general. What do you think the creation side is of this?

19:26Obviously, there are these, you know, image models and video models that people are starting to experiment with. But what's sort of the mature view of this? And where does this go in, you know, two to three years for how we all create media? Yeah, I think it's a fascinating question. And I think you can look at two different types of services. So Spotify, actually, both in music, podcast, and books, we're an aggregation service, right? So we want to aggregate as many creators as possible to have the biggest possible catalog for our consumers. And then our task is to understand what of this vast catalog the consumer likes and recommend that just lowering friction we didn't innovate music we didn't innovate podcasts and we didn't innovate books but i think what's interesting about if you look at something like tiktok that we mentioned is they did a great innovation on the consumption side by this full screen ui that gives like perfect feedback to an explore exploit algorithm but what i think a lot of people miss unless you use tiktok a lot as a as a creator is they also did a lot of innovation on the creator side right so because you create in tiktok they actually use a ton of AI to drastically lower the friction on the creation side.

20:33And they needed to do that because they were not programming an existing format like we do. They were starting a new format, whatever you wanted to call that music sync dance video that originally came from Musical.ly. So actually, they did a lot of innovation on both sides. And I think the AI innovation on the creator side was probably as important as the AI innovation on the consumer side. Can you give some examples? Like, what are some of the things they did on the creation side? Well, for example, when you create a talk, you know, it can synchronize the music, help you synchronize moves, all of these things.

21:09It's actually very, and it's increasingly very, very AI driven. And so I think that if you're building a new service now that you want to compete with YouTube and TikTok and Instagram and so forth, that is the vector. Use AI on the creation side somehow. to create a new type of format or interaction that drastically changes the cost curve or the friction of creating content. So what I think is going to happen for us is something more traditional, namely that our goal is not to replace the creators, it is to get more creators, to make them more productive, right? That is what maximizes the value of Spotify for consumers, to have more music, more podcasts, more audiobooks.

21:52And I certainly think AI will increase the productivity of musicians, podcasters, and authors. Do you imagine a world in which we talked about TikTok and we talked about how they're already using AI to help make the creation easier to establish their own format? If you just stick with that trajectory, do you imagine a world in which a platform like TikTok is generating the content themselves and delivering the perfect piece of content at the perfect time? I mean, it's a great question. So I think if you want to start a new company and compete with TikTok, that would be the dream, you know, sort of infinite content at no marginal cost and just generate and generate.

22:33And you can imagine like a future Netflix where it just generates every possible movie that could be generated and renders it, right? So I can't say that is theoretically impossible. It doesn't seem impossible. I don't think it's very likely, though. I think what would happen is more what happened with TikTok, which means that you actually lower the friction to human creation. You make many more people much more powerful and productive. It still seems like you need the human idea. I think an example is, if you look at what happened with text generation and audio, it is already today fully possible to generate a full algorithmic podcast.

23:15Just prompt LLM to have an interesting discussion and then render it. And in fact, it's being done. It's uploaded. I can tell you they're just not very interesting. I can't say exactly why. It could be that it's just like a few more iterations and they will be amazing. But they're not yet. But I can't argue theoretically why it wouldn't be possible. I just don't see it yet, even though the quality is there. And I don't predict that it will happen anytime soon. I wonder if it's because, you know, these models are trained on all of the media and content that exists in the world today. And so anything that gets spit out almost, it feels like there's a ceiling, there's an upper limit, what's already been created.

24:00but the greatest works of art always feel like they're reaching new heights, right? People do something innovative or different that you haven't heard before, and it raises the bar. These models are sort of hitting the current bar. I wonder if that has something to do with it. I think it's a great point, and that may be it, that these models are specifically trained to predict and repeat what has already been done with some variation, whereas what we like are things that never happened before, and it would be harder for these models to do that. I think a machine learning scientist would argue like, that's not really true.

24:31You could have randomization. So I can't say it's impossible. I can just say it doesn't seem like the case yet. Yeah. So speaking of that, you know, all these models today, they're being trained on the public internet. They're crawling anything that's out there. Oftentimes things that are copyright protected. It feels like training data and training models is the wild west right now. There's just no rules. People are doing whatever they want. Does it stay that way? What is the future of training data for language models, video models, image models, music models, all these things? What does this look like in the future?

25:06So I obviously don't have a crystal ball. And I think there are a couple of ways to think about it. One is, what do I think will happen in the world? And then also, what is Spotify's view, regardless of what happens in the world? and I think there is going to be legislation at the end of the day and people are going to have to adapt to legislation that will be the lower bar and the legislation could be that turns out you know it's legal to train on anything and then I think it's a market economy some companies will do that and then to compete other companies will and then that becomes the status quo it could be that legislation says that that's not the case that it emerges that you know you need to somehow reimburse or respect people who don't want their data to participate, right?

25:53So that will be the sort of the lower bar is what legislation sort of dictates. And I think it's very possible that legislation will dictate a reimbursement model. You know, you can compare to the Wild West of piracy that existed for a while. It doesn't really exist anymore. So I don't think the world has to be in that Wild West stage. I think some order will emerge. From Spotify's point of view, as I said, our view is to have as many artists as possible on Spotify. And we actually want the artists to create more music, not to replace them. So from our point of view, we certainly respect their data.

26:33And regardless of if there are like loopholes that we could and so forth. And I think that's why Spotify I started was because the music model worked for consumers, but not for creators. So I think it'd be a step back to create a model that again, works for consumers, but not for creators. So certainly for us, the goal is to find a model where you can leverage the new technology and it works for the creator ecosystem. But obviously when disruptions happen, as with piracy, there is often a period of time first where it only works for one side of the marketplace. But that's not long-term sustainable.

Read the full transcript

27:16Yeah, it seems like there could be legislation, but it almost seems impossible to build a system that could actually pull this off. So I don't know. I just wonder how that would play out. Maybe that's where crypto comes back. I was just going to say, yeah, some people like Fred Wilson have recently said that, you know, this could be the application for the blockchain, right? So everything, every piece of content that gets created basically gets minted to the blockchain and you can basically trace the provenance of a piece of content all the way back. I actually think it's an interesting use case.

27:51Yeah, I don't think he's wrong. I think it's potentially, I mean, for AI in general. So one of the problems you say, like, it's so complicated, right? Already. Because there's the artists, there's the songwriters, mechanical rights, performing rights. You have so many different people involved. But it's only complex because you have humans doing it. If that was on a blockchain, you can make a billion transactions per day. it's actually not, it's just computation. So I think if you instrumented it programmatically, you could solve it. It just looks very complicated to humans. So I think he's right in that.

28:32And in general, I think the problem of authenticity, traceability, and who is who, what is the fake content versus not, it is potentially one of the big applications for the blockchain. I agree with him on that. And it also seems like it would be made so much harder because there's so much content that exists today that was probably inspired by some other work that we don't have record of anywhere, right? I mean, this almost maps to the existing royalty and infrastructure we have in music today, which is so, so messy, right? And it's across so many different systems. I don't know. This almost sounds like an unsolvable problem, but I'm sure some smart people will figure it out.

29:14But let's talk about ChatGPT and sort of AI's impact on kind of product and business. ChatGPT was obviously the biggest story of the year last year, you know, reported to have in the hundreds of millions of users generating north of a billion dollars annually in revenue. It feels like - Pretty impressive first year. Yeah, crazy. You know, and so everyone talks about it like ChatGPT is sort of the iPhone of AI. It's like it was the iPhone moment. But, you know, something that strikes me about this is it's a text conversation, right? You're chatting with an agent. And text as a UI seems like an outdated primitive, right?

29:56It feels like sort of a weird primitive to build a whole new platform and interface on top of. What do you think of that? So I agree. I think the premise that people use is you see it with, you know, the demos of, for example, the new Rabbit hardware, the paradigm that everyone is looking for is intelligence replaces apps or replaces separate UIs. And where does that come from? Well, if you have an executive assistant, for example, then you can say that you want something achieved and that person can work across several services and UIs to achieve that task, right? So I think it's a very rational thought that, you know, intelligence would replace the need for several different apps.

30:46And I think the best analogy is to anthropomorphize it and say, like, if you had a really smart person helping you or working for you, what could you achieve? And then you could assume that AI could probably achieve that if it just keeps scaling. So that seems pretty reasonable to me. But I also think that people take it too far. and and you know anthropomorphizing has its benefits because you can predict the future like okay so if i can do this with a very smart friend now and you think ai is going to get that smart in a year then okay i can build that product in a year so it's helpful but the problem with anthropomorphizing is that it can you can only predict what a human could have done not what an ai that can also read images or it's much faster than a human could have done or or, you know, that can also render images or, you know, generate it.

31:38So it limits you. I think it's both useful to anthropomorphize in the near term. I think it's dangerous to anthropomorphize and think it can only do what a human, what a single human could have done in the long term, because it can probably do much more and things that a human could never do. Right. So I think an interesting example is, and the framework that I use is, if you take that path of like, you know, I need to use five apps to complete this task. Now I have this UI, this AI that I can talk to through text or voice, and it completes them, whether it uses API or it uses sort of these, you know, it interacts directly with the UI as, for example, the rabbit action model.

32:19It doesn't really matter. I think that seems very likely to work for productivity tasks, where there is something you didn't want to do. That gets a lot better, right? And it works with the anthropomorphization. Like you ask someone else to do it because you didn't want to do it. But I don't think it's that helpful when you think about entertainment. It makes no sense for me to ask my friend, like, could you just go and watch Netflix for me for a while? Because I really don't have time. Can you go and listen to Spotify? Like the whole point with entertainment, I think it works where you want to save time.

32:55Productivity is about saving time. Entertainment is about wasting time. So I don't think it is a very good framework for entertainment services. And so an example would be, if you look at something like Spotify, of course, you could say like, hey, Spotify, play this song. But that's a very narrow use case. And you can already do that with like a Google Home or something. but if you imagine that you had to say like hey ai can you go and scroll through the front page and tell me what you see so that i can choose something then a visual ui is going to kick ass it's going to be so much more effective because it's two-dimensional you have images right so i do not think that text will replace visual uis that doesn't make any sense to me i think that's taken it too far i still think like a user interface that can present images moving pictures and actually play sound is going to be vastly better than an AI in between that tries to relate what you would have seen if you looked at the UI, right?

33:55So personally, I don't think a Netflix or Spotify or Instagram is going to sort of disappear into the background and be disintermediated by a text box. It doesn't really make sense from a productivity point of view because you're not trying to save time. You're actually trying to waste time. what I think is exciting though so if you don't anthropomorphize and you think instead what could a maximum product be not what could a human have done if you talked to them over the phone but if you said like which is a helpful model for some productivity tasks but not for entertainment I think you know the promise is an AI that can understand your intent when you speak and other inputs if you have them but I certainly think voice is a text is a very strong input, but can also render user interfaces that are much more dynamic.

34:46Today, user interfaces have to be very specific and repetitive because they're not generated on demand, right? They're pre-programmed. So you have to think through which views you want in an app and so forth. You could imagine far into the future because these things can generate code. If you squint a little bit, you could almost imagine that you're simulating an app like Spotify and it's literally like rendering the app or the code in real time. And I don't think... Totally dynamic UI. Yeah, I don't think it will go that far, but you could go some ways. You could have like the search view could start becoming more dynamic.

35:23And sometimes it renders images if that's helpful for the search result. Sometimes it renders text fields. So you could imagine that the whole thing gets more dynamic and intelligent. Do you think the starting point for applications though become either text to voice? Like, is that now the new home screen, right? You go to your computer, your phone, and it's just a prompt, right? And that sort of starts everything or whatever you're going to be doing, be it wasting time or spending time or saving time. No, I don't think so. Because, again, a prompt is only helpful if you know exactly what you want to do.

35:58If you have a task in mind that you want to do faster, an input box is great. but but you know then then if you take that if you think that's enough then Spotify should just be the search box on the front page right and we know we've tried if you if you put the search box there people struggle what to listen to right people don't know what they want to do all the time so I don't think it will be only that but I do think that will be just a search is a big big part of any service it will be a big part of your life for many tasks and and back to like anthropomorphizing think through like if you had a super highly intelligent very effective friend you know that worked for you for free what would you use them for versus what wouldn't you use them for you know people are talking about ai as a platform shift and trying to compare it to previous platform shifts oh it's the new iphone oh it's the new cloud computing i tend to think of it as a resource like capital or labor or work right um do you see it the same way and sort of How do you think it'll impact the business model of software moving forward?

37:03So I do see it sort of the same way. I think it's good to try to think about something different. I think one way to think about AI is that, you know, we have mechanical labor and that has been automated. And what's happening now is that sort of cognitive labor is being automated. and it's a good framework to think through you know what happened when mechanical labor got automated and then try to figure out what might happen when cognitive labor gets automated so it's like cognitive machines in addition to sort of mechanical machines i think thinking of it as as literally intelligence like units of intelligence is interesting so you know if you it sounds weird like what do you mean with units of intelligence and like you know you're buying units of intelligence on tap that makes no sense and then you think about your hr department and what they're doing and they're hiring like units of intelligence every day that's what you're doing you're trying to get as much intelligence into your company as possible right but you're paying a lot for the best intelligence you have all these tests you know to figure out you know and And at the core, the way you try to compete as a company is to have the most intelligence in-house.

38:22People talk about talent, density, and so forth, right? And so if you think about it like that, it is intelligence, but now you can buy it sort of on tap or per unit. Then I think that can be helpful. Because the tricky thing, to your point, is it is so general. Intelligence is by definition completely general. They almost need to think of it as some sort of resource. And so I literally think of it as intelligence. But you get intelligence and you get execution, right, in a sense, for certain tasks. So it really is, the HR department hiring people is a really helpful analogy. It's scary to think just how massive it could be in this new framework, in this new model.

39:12I think what's going to happen sort of to sort of preempt the question maybe of labor and so forth, because you could say like, oh, now you can buy intelligence. The second that's one cent cheaper than hiring intelligence, you're only going to use computers. I don't think that's going to happen. It's the same as with musicians. That's not what we see today. What we do see very clearly is like a developer with co-pilot is more productive than a developer without. out. So if you can buy more intelligence for that developer, you know, if you do a Ray Kurzweil and say like, hey, the nanobots are here, you cannot buy neocortex in the cloud.

39:46Would you as a company buy more neocortex for your developers? Yes, you would. And I think buying Copilot is like a weak analogy of buying more neocortex for your developer. And so I think companies are going to start spending more and more on their existing staff to make them more and more productive. But I don't think in this competitive economy economy you're actually going to reduce your labor force because then someone else is going to take their their opex and and uh out compete you so i think the pressure is to make your existing workforce more and more productive speaking of intelligence uh gbt5 is on the horizon um what what do you think the impact of this is going to be how much how much more intelligent do you think it's going to be or feel than what we experience today?

40:37I mean, I don't know. We've lived through a couple of these now with GPT-2, GPT-3, 3.5, with RLHF and then GPT-4. So to my previous point, in a way, I think it's pretty predictable that you will be blown away. So you won't be as blown away because you sat there and expected to be blown away it's this weird thing when someone tells you a move is amazing you're not as impressed as if you had no expectations so personally i'm expected expecting to be blown away and i probably will be but that actually means it's more predictable what i think is i mean i don't know that much or actually i know almost nothing at all about gpt5 i only know what most people know my expectation is that it is more predictable because it's now going to get predictably better at existing dimensions what happened at previous times was that you had these emerging capabilities where it did completely new things right and that is what surprised us you know when you have like chain of thought you're like jesus this thing is reasoning right and it couldn't do math at all like not at all couldn't do one plus two and then like oh it can do math now you can reason around math and theory.

41:48So I don't think we'll get as surprised. I'm hoping I'm wrong here. Be interesting. But I don't think we'll be as surprised about completely new capabilities. We'll be blown away by how good the existing capabilities got. You know, they will be way past most humans. And I think you see this. Why am I saying this? Well, because if you look at the tests that are out there, these models now are pretty good at almost all human capabilities. So I don't actually you know, what it would be that emerged that it can't do at all anymore. It can sort of do everything. But if you compare it to image recognition, there was this, you know, forever image recognition didn't get any better.

42:29Like it barely couldn't recognize images at all. It was close to random before sort of deep learning came along. And then this race started and it got better than like random. and then you've got 60, 70, 80, 90, 95%, 96, 97%. And actually humans are only like something like 96, 97 % and it beat humans. And I think we're already on that ramp up trajectory of the dimensions itself, whether it's reasoning around physics and math or taking SATs or talking about legislation or something like that. So I would be, what would surprise me is some completely new ability that we hadn't seen at all. What I'm expecting is to be blown away by how good the existing capabilities now are.

43:18And, you know, hallucinations will be very few and far between. The reasoning capabilities will be much stronger. You can probably, you know, reason for long, spend many tokens on getting to deeper reasoning. What would be really cool, which is very speculative, I think, is I think a lot of people in AI in general, for obvious reasons, are very interested in math. And I think theoretical math is the ultimate frontier. Can you create new math proofs that didn't exist before? So maybe one of the abilities, if I'm going to correct myself, that isn't really proven yet is can you actually make something new that didn't exist?

44:02And some people would say, of course you can. you just raise the temperature a bit and then you get some text that never existed. This is new. Other people would say like, it's not really new. You know, I want some sort of deductive reasoning, like a law of physics, not just something that was very close to the law of physics or something that existed. And I think what is interesting now is, you know, these models are starting to get this ability to reason at a higher level. Some people claim it's not reasoning. It's just reasoning for a while, or if you pretend that it is reasoning for a while, then you have this emergent space on top of just the pattern recognition that is happening.

44:46And you could imagine sort of doing something like reinforcement learning on that space where you ask it to reason through hundreds of thousands or millions of possible math theorems based on everything you learn about math. But the beauty of math is you can then test if the theorem was true. It's one of these one-way functions that is hard to come up with, easy to prove if it's true. It's like crypto, right? It's very hard to break it, but you can very quickly verify it if it's true. So then you could imagine this agent that using an LLM tries to reason through, just brute force reason through all possible math theorems, sort of randomly like a reinforcement learning process.

45:35But it can verify. And it's going to reason based on pattern recognition, like it's seen this type of reasoning before. And maybe it could come up with math proofs that just never existed before. And then you could verify offline that they were true theorems. That might happen. I think an interesting analogy here is this mathematician named Ramanuyan, Indian mathematician, who is a great movie capturing this. But basically, he grew up very poorly in India. So self-taught mathematician just on the streets of, I don't remember which city in India. But he somehow developed this incredibly strong instinct for math.

46:17and it was discovered by, his math was discovered by some British mathematicians and he was sort of brought to England. But he had never learned proper math. So his theorems came to him in dreams from a god. And so when he was asked, like, how do you know this theorem is true? He answered, well, of course it's true. Like a god told me, right? And it turned out that actually many of his theorems were true, but not all of them. And then he learned formal math and he could figure out how to prove which of his dreams were true. So you could imagine AI, and what is cool about Ramanujan is, I think most mathematicians would say that math is not pattern recognition.

46:58It is true logical reasoning. But then you can't explain Ramanujan. He didn't have the logical reasoning, the form of math. He just had dreams. So maybe even math is sort of exploring exploration of pattern recognition. But then you need this mechanism of form of math to verify which is true. So that's a long-winded answer. Maybe math will be the thing that surprises us with GPT-5. OpenAI keeps talking about how AGI is coming. You know, we're getting closer and closer. Maybe that'll be GPT-5. But at this point, like, what is AGI? And does the definition even matter? I mean, you mentioned, you know, the models are reasoning.

47:39Are they reasoning or does it just seem like reasoning? And like I said, does it even matter? If we imagine that it's reasoning, isn't it kind of the same thing as it is reasoning? So I actually think you're right in that it doesn't matter. I think if you ask about the definition of AGI or even just I, intelligence, I think the best definition I've heard is probably Marcus Hutters of Hutter Prize fame. And he says that intelligence is the ability to achieve complex goals in a wide variety of environments. which sounds like a mouthful, but it's actually pretty straightforward. It's just the ability to achieve goals, complex goals in different settings.

48:25And humans are at some level on that spectrum. And actually, we're not at the same level. Einstein was quite a bit sharper than the rest of us, right? So it's not even a level to your point of like, does it really matter? So I don't think it matters. I think we've had super intelligence among us, you know, the Einsteins of the world. And that is also pretty promising because we seem to be doing fine. You know, he came up with relativity and you can argue that new innovations, you know, like nuclear and so forth came out of that. And maybe that will ruin us someday. But certainly everyone doesn't have the same intelligence.

49:02We're on the spectrum there. And that seems to work. It's unclear what happens if you go really, really, really, really far on that spectrum. I can't say that I know. But for that reason, I agree with you. I don't think just passing like the human level really matters because some humans already passed the average human level. Achieving goals in a wide variety of environments is interesting. It makes me think of these products like, well, products like the rabbit and more broadly, this notion of an action model, right? Everything right now, we're talking about language models, the rabbit and, you know, people that are thinking about that type of innovation are talking about action.

49:42models when do you think we're going to make the leap from generation of content and and writing and images and all this stuff to action taking and and what will be sort of the nearest term implications of this well i think you know to your point of of the rabbit and so forth like it's already happening just very early so you know one answer will be next week maybe who knows but but But my point is, I think it's quite imminent. I think someone more skeptical would say like, well, we've had the idea of agents for a good while now and they're really cool, but they're not that useful yet. I think that's also fair.

50:23But it seems pretty straightforward that at least, you know, semi-specialized agents that help you complete like something that would have required two tasks or three tasks or four tasks for specific use cases like traveling. Like, I don't see how it couldn't happen. So I think these agents are sort of going to sneak up on us. And I'm not so sure you're going to realize that you're talking to an agent. It might just look like when you used to talk to three services, you're now just talking to one and it's three behind the scenes. Right. And then, you know, do we get very general agents that does everything for you?

50:59Or is it going to be that, you know, each sort of service has its own agent? I don't know. But I think like action transformers and taking actions is going to happen very soon. I know that inside companies, I think Amazon published a paper this fall where they are talking about sort of these orchestrator LLMs, you know, from having worked at Spotify that a backend in machine learning is usually like, it's not the algorithm. It's actually like hundreds of different mini algorithms, right? They're all like individually tuned and so forth. and then together they sort of produce something that is almost like alchemy.

51:42You don't know how all these different algorithms are going to interact with each other. So you try to tune and understand. But the result looks like an algorithm. It's actually many algorithms. I think that is changing. I think people are starting to put these sort of orchestrator LLMs in front of their system that actually is more akin to like a single brain or algorithm that does take all the signals from you. and understands that, looks at the embeddings of your user history, for example, and then tries to talk to your APIs. So there's a layer in between the user and your APIs, which is an LLM that says like, oh, Mike is now saying this to his Alexa or typing this into Spotify or something.

52:21And then the LLM, instead of hard-coded rules for what will happen now, call this API and that API, the LLM says like, well, based on what Mike said and reasoning through chain of thought, I should probably call this API and do this for Mike and then this and then send this back to the screen. And that is an agent. You don't see it. You're not going to understand it's an agent, but it is an agent inside something like Spotify that reasons through what you did in the UI and actually talks to the backend. So I can see a future where a company builds like lots of capabilities, lots of APIs internally, and then just has a really big LLM fronting the user.

52:56And the LLM actually reasons through which APIs to call and what to do for the user, which is pretty cool because it could then do things that the programmer didn't necessarily predict, if you understand what I mean. And I think certainly if you want to build a voice assistant, that seems much more scalable than building these one-off custom flows for every possible use case that you could imagine. So I think agents are coming in various forms, and they're sort of already here in some forms, even though you don't see them. Speaking of agents and taking actions and AGI, I think about this analogy that Ilya from OpenAI gave about the risk of AI, comparing it to the threat of humans to animals, right?

53:41When humans have a goal that indirectly impacts animals, humans often don't stop to think about that impact. How do you think about AI and agents taking actions when the threat or the risk of the actions they take impact humans? yeah so i would just start out by saying i'm not particularly concerned about the existential risk that many people are for various reasons and you know you probably shouldn't listen to me because i'm not really an ai expert but i am more concerned about the practical risks in the near term okay um you know it's like when you introduce anything new whether it's like uh you know cigarettes but that's a bad analogy because it's only bad like cars was going to have some negative effect on people's health because they stopped walking right i think there will be consequences i think it's dumb to be naive and say that this is the one technology which has no risks and will have no consequences of course there will have consequences in their risks so first of all i want to say i think it's very reasonable to be careful and invest in safety and back to predicting prediction don't just try to predict the good things that could come which are many, like AlphaFold and solving cancer and so forth, I think that will happen.

54:56But also try to predict the bad things and try to prevent them. That seems very reasonable to me. But on the timescale, I'm more worried about the near-term DOM AIs than the potential risk of the very smart future AIs. And so one question to ask yourself is, do you think the problem with the world today is that we have too much intelligence or too little? and if you ask me that question you know are you most worried about the most intelligent people around you the least intelligent people around you which cause the most havoc and uh if you think about you know ilia's point about who cares about other species is it that you know it seems like it's the smartest among us are the ones who seem to sort of care the most for other species because they understand that actually if if that bee over there goes extinct eventually that's my environment and my climate if you're smart and really understand causality or at least correlation very deeply you're going to get more careful it's actually when you're not that smart that you're dangerous right so from a very high level i would think that more intelligence is a good thing in the world the problem isn't that we have too much intelligence i think so one maybe the problem is to get these ais smart enough to start caring about and understanding its full ecosystem including the humans that actually are building and powering it.

56:20It doesn't mean we're smart from an AI to make humans go extinct. Certainly not too soon, right? So if you took in the very far future where somehow the AIs are fully self-sufficient, they run and build somehow all the factories with robots, then maybe. But I think that's extrapolating too far. In the foreseeable future, it'll be very dumb for the AI. It's only if it doesn't understand enough that it's going to make humans extinct. So maybe the big risk are the stupid AIs. and i don't know about this is very speculative but you know humans created empathy as sort of a feeling and you can argue that that's sort of some sort of divine good but if you're if you're not sort of creationist it has to have emerged from evolution somehow there must be some benefit to empathy and i think that since these you know since these things are trained on our thoughts and the task is to predict what we you know what we think what comes next and empathy seems to be a very important part of how we reason if you read all the text on the internet and you try to predict the next token if you want to predict how we think and how we come to our conclusions in these sentences it doesn't seem unlikely to me that these models would would develop or at least emulate something like empathy in order to achieve its goals right And I think this is an important point I want to make, that while people like Max Tegmark and others, they often call this an alien intelligence, and it scares people.

57:57It's like the aliens already landed on Earth. They're here now. It's called AI. I don't think that's the right way to think about it. I actually think it's the opposite. This intelligence is literally physically modeled on our brains and neurons. We really looked at the brain and, you know, first and foremost, the eyes. And we modeled the artificial neuron after the biological. So physically, it's modeled after us. And then we trained it on all our thoughts, exactly how we reason. So this is probably by far the most human intelligence we'll ever encounter because it is physically built. It works physically like human intelligence, and it is also trained on human intelligence.

58:41So, you know, I think calling this alien intelligence is the wrong way to think about it. And alien intelligence would be something vastly different. Everything we see is that these things get more and more like us. So if it's smart enough and it tries to model us, hopefully we'll understand its ecosystem, just as we are beginning to do now that we're getting smarter as a society, and we're trying to correct our ways instead of eradicating ourselves. So that's how I think about it. Gustav, this is a fascinating way to end the conversation and frankly, an optimistic one. Thank you so much for the time today.

59:20You are a very busy person. So we really appreciate you coming on. Thanks for having me. This was a great conversation. Thanks so much for listening to Generative Now. If you liked what you heard, please do us a favor and rate and review the podcast on Spotify and Apple Podcasts. And if you want to learn more, follow at Lightspeed on X, LinkedIn, Instagram, and everywhere else. Jennerd now is produced by Lightspeed in partnership with Pod People. We will be back next week. See you then.

From the publisher

Spotify’s personalized recommendations have set it apart from the pack of music streaming platforms for years. Those curated recommendations have only gotten stronger with the onset of AI, with the ability to power things like your own personal AI DJ. This week on the podcast, we’re revisiting a conversation with Spotify’s Co-President, CPO & CTO, Gustav Söderström. Gustav joins host and Lightspeed Partner Michael Mignano to talk about building UX that can power more efficient AI, the future of AGI, and what the technology could look live over the next few iterations.  


Episode Chapters

(00:00) An introduction to Gustav Söderström

(01:29) How Spotify became one of the first consumer AI companies

(05:57) A philosophical shift: UI that works for AI, not the other way around

(12:52) “Talk is cheap, so do a lot of it”

(15:00) Building for the tech we will have, not the tech we do have

(19:43) What will AI do for media creation?

(25:00) The Wild West of model training

(29:45) How anthropomorphizing AI can help you understand its potential

(37:04) What happens when cognitive labor gets automated?

(40:42) How much more intelligent could AI get?

(47:45) Is AGI on the horizon, or already here?

(53:48) Should we be worried about AGI?


Stay in touch:

The content here does not constitute tax, legal, business or investment advice or an offer to provide such advice, should not be construed as advocating the purchase or sale of any security or investment or a recommendation of any company, and is not an offer, or solicitation of an offer, for the purchase or sale of any security or investment product. For more details please see lsvp.com/legal.

More from Generative Now | AI Builders on Creating the Future

All 90 episodes
Gustav Söderström: Catapulting Spotify to the Front of the AI Revolution (Encore)Generative Now | AI Builders on Creating the Future · 1 h
Listen in VO