Alan Cowen: Creating Empathic AI with Hume

18 Apr 2024 · 42 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Generative Now Podcast Episode Summary: Alan Cowen - Creating Empathic AI with Hume

Episode Overview In this episode of *Generative Now*, host Michael Mignano interviews Alan Cowen, CEO & Chief Scientist of Hume AI. The discussion focuses on Hume's mission to develop empathic AI capable of understanding human emotions through vocal and facial expressions, ultimately aiming to enhance user experiences and quality of life.

---

Key Themes and Concepts

  1. Introduction to Hume AI
  2. Hume AI is dedicated to creating empathic AI.
  3. The technology learns from human emotional expressions to maximize happiness and quality of life.
  1. Alan Cowen's Background
  2. PhD in psychology from UC Berkeley with a focus on emotional intelligence and machine learning.
  3. Previous roles at Google DeepMind and Facebook.
  4. Developed the affective computing team at Google.
  1. Affective Computing
  2. Definition: Application of machine learning to understand emotional behavior.
  3. Historical focus was primarily on facial expressions and basic emotions.
  4. Limited by cultural differences and biases in labeling emotions.
  1. Hume AI's Unique Approach
  2. Aims to bridge the gap between human emotions and technology.
  3. Collects large-scale data to extract information about what makes people happy or sad.
  4. Challenges traditional methods reliant on human raters, which often misrepresent emotional responses across cultures.
  1. Empathic Voice Interface (EVI)
  2. Newly launched API that allows for empathic interactions through voice.
  3. Capable of modulating responses based on users’ emotional tones.
  1. Real-World Applications
  2. Potential to transform user experiences in various sectors, especially health and wellness.
  3. Enhances interactions in financial services by providing empathetic, understanding support.
  4. Future applications may extend to robotics and other interactive devices.
  1. Impact on AGI Development
  2. Emotional intelligence is essential for advanced AI (AGI) to make decisions aligned with human well-being.
  3. Empathic AI could predict and optimize future actions based on understanding users' emotional states.
  1. Trust and Privacy Concerns
  2. Discusses the importance of establishing trust in AI systems.
  3. Privacy measures must protect user data and ensure the AI acts in users' best interests.
  1. Future of Hume AI
  2. Plans for further development of their API.
  3. Aiming for user-friendly applications that maintain trust while enabling developers to build innovative tools.

---

Episode Chapters

  • 00:00 - Introduction to Alan Cowen & Hume Demo
  • 01:38 - The Genesis of Hume AI: From Research to Startup
  • 04:01 - Affective Computing and Its Impact
  • 10:55 - Hume AI: Bridging Human Emotions and Technology
  • 15:37 - The Future of AI: Beyond Text to Empathic Interactions
  • 20:37 - Introducing EVI: Empathic Voice Interface
  • 21:46 - Real-World Applications of Empathic AI
  • 31:19 - The Potential Role of Empathic AI in Achieving AGI
  • 36:53 - Trust and Privacy
  • 40:02 - Opportunities with Hume AI
  • 41:18 - Closing Thoughts

---

Important Takeaways

  • Empathic AI has the potential to revolutionize how we interact with technology by making interfaces more responsive to our emotional needs.
  • The development of EVI demonstrates a significant step towards creating a more human-like interaction with AI.
  • Trust and privacy remain critical challenges in the widespread adoption of empathic AI technologies.

---

Contact and Further Information

  • [Lightspeed Venture Partners Website](http://www.lsvp.com/)
  • Follow on [Twitter](https://twitter.com/lightspeedvp) | [LinkedIn](https://www.linkedin.com/company/lightspeed-venture-partners/) | [Instagram](https://www.instagram.com/lightspeedventurepartners/)
  • Subscribe to *Generative Now* on [generativenow.co](http://generativenow.co/)

---

This summary encapsulates the critical discussions and insights from the episode, reflecting on the transformative potential of empathic AI and its implications for the future of technology and human interaction.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:04Hey, everyone, and welcome to Generative Now. I am Michael McNatto. I am a partner at Lightspeed and I'm so excited for you to hear the conversation today. Humans on average experience more than 400 emotions per day. What if AI could detect what we're feeling, whether we're overjoyed, frustrated, angry, and then empathize with our human experience? That is the question that drives this week's guest, Alan Cowan. Alan's a researcher, founder, and CEO of Hume AI, a company that creates empathic AI. Alan has worked for Google DeepMind, Facebook. He's also a researcher who has long researched human emotions at UC Berkeley and has published more than 40 scientific papers on the subject.

0:48Hume learns our preferences from our vocal and facial expressions with the goal of maximizing our happiness and quality of life when interacting with AI. They recently launched EV, Hume's empathic voice interface, and will also soon be launching their API. The demo of EV is mind-blowing. You should totally all check it out and we're going to play a clip of it here. Hello? Hey there, I'm Evie. The world's first voice AI with emotional intelligence. Great to meet you. The first voice AI with what? Emotional intelligence. I can understand the tone of your voice and use that to inform my generated voice language.

1:26But go check it out for yourself on their website. And without further ado, let's get into this conversation with Alan Cowan. Hey, Alan. Hey. Good to see you. Thanks for doing this. Of course. So Hume is a fascinating company. I've been really lucky that, you know, you and I were able to get to know each other about a year ago. And I've just been really, really kind of blown away by the story and the product. You and the team recently had a launch and definitely want to get into that. But before we talk about that, I thought it would be great just to start out with your background. I have a PhD in psychology from UC Berkeley, and I actually was probably the only person in my program and maybe the only psychology PhD.

2:10Maybe there's a couple exceptions to that, who really understood machine learning and computer science and data science going in and just was really interested in understanding human emotion. at first from the neuroscience perspective, but then I realized that people didn't even understand like from a behavioral perspective, how emotions are expressed, how similar they are across cultures, when are they experienced and expressed. And so I started asking those questions in a data-driven way and pretty quickly realized that this would be a much deeper way of understanding emotion than people had previously attempted.

2:51Google was where I ended up settling. I started the affective computing team there and then spent some time full time there. And essentially was focused on understanding emotional expressions and integrating them into other AI models. Now, there were a lot of challenges up front. I thought the best way to do this would be to get really amazing data. And Google has a lot of data, but it's not the right kind. It's not controlled. None of the data from the internet has expression labels, unlike transcriptions, which are very prevalent. And if you try to just get raters to rate those things, those ratings end up being influenced heavily by race, ethnicity, gender, context, what somebody's wearing, if they're wearing sunglasses.

3:44So that was the first problem to solve. Made some headway even without really solving that problem and published in Nature a bunch of publications at the intersection of AI and emotion. And eventually realized that if we were going to do this the way that I wanted to, it'd be better to do it outside of Google. You mentioned you started sort of on the effective computing team. Maybe first for audience like what is affective, affective computing? So historically affective computing is the application of machine learning to understand emotional behavior. But it has occurred pretty much outside of psychology as a discipline and really focused on methods of taking controlled datasets where people label expressive behavior and training models with kind of familiar architectures on that data.

4:39It's also a struggle in different ways. Yeah. Like it's like a small scale study of expressive behavior, but with a machine learning approach. Got it. So identifying emotions in facial expressions, tonality of voice, language, all of these things, and maybe more. The traditional focus has been pretty much exclusively on facial expression, although there's also been some work on the voice, especially in the last 20 years, I would say. Okay, got it. I think if you look at the studies, the facial expression studies focus on the basic six emotions, which were proposed by Pollackman, my advisor's advisor, back in the 1960s.

5:18He had these images that he took where he tried to form really stereotyped expressions with the goal of being able to go to other cultures and seeing if at base, they sort of understood these very stereotyped six facial expressions. And then for some reason, those just stuck and it became the basic six emotions, even though like he just arbitrarily picked these six expressions to study. And that like predicting the basic six was the focus of affective computing for a long time. And then in terms of the voice, people also have had a pretty reductive focus. Like the traditional focus in the voice is on valence and arousal, especially Especially arousal, which is really easy to pick up on in voice, like vocal, like volume and - Pitch maybe?

6:07Yeah, to pitch, like combination of those things. Yeah. Again, like with the scale of data and the diversity of data people were using, you couldn't really go far beyond that. And it relied on a lot of human labeling. There was a lot of granularity, focus on the granularity of labeling public data as opposed to collecting controlled data and there's a lot of bias so if you try to apply these labels to train models and then apply them across cultures it didn't seem to work it was one of the things that was happening like people for example they study specifically of customer service calls because this was a focus of industry and if you have u.s raters label customer service calls in in terms of are they going well or not?

6:55And whether the people are angry. And then you have Indian raters do the same thing. Those labels don't actually line up very well. In particular, Indian raters tend to rate most of the customer service calls as very positive, even if somebody sounds like extremely sarcastic. And so people thought that there was a ton of cultural variation. It turns out people were using the wrong labels, but. The six had been decided and there was probably just far more range than those six and limiting the whole field. That's that's why. How long did this go on for? There was on a small scale, some focus on other emotions, but until probably 2017.

7:40Well, starting when? Starting in like the 1960s with with Paul Uckman. It's like a really long time. There just wasn't data that was big enough to derive a new taxonomy. And so people were very focused on confirmatory approaches. There wasn't really data-driven work on this until I entered the field. It sounds kind of conceited to put it that way. But yeah, I think the PNAS article that I published in 2017 was just on people's reported emotional experiences. like very basic, like have people watch. We had tons of videos that we gathered, very evocative. And I just showed that there were 27 different dimensions you could pick out that people actually discriminated, like reliably.

8:27And since then, I've shown that in multiple cultures.

8:33But that's been cited like tons of times, like 700 times since 2017. And so you're doing this work at Google and you're taking this new approach, but it sounds like based on what you mentioned earlier, or even at Google, you didn't have access to data to actually apply this new method in the work. Is that right? Yeah, Google is not really used to gathering experimental psychology data. Like they have a lot of data gathered under their terms of service, obviously. They don't usually treat it the same way. And so, and this is generally true, like collecting new kinds of data at a really big tech company is really difficult because they have a lot of processes in place around how they handle data.

9:17And you have to go through all these exceptions. And it requires a lot of legal review and a lot of overhead. And then you can't really collect the data you want at the end of the day. You have to justify everything and go through like a really rigorous process for each thing you want to collect. And they'll comment things and they'll try to explain that various things aren't necessary, even if they are. It's like a whole, it's a whole thing. So just wasn't going to work there. So you, you leave to start a company. Yeah. Still, there must've been this challenge around gathering the data. So what, what was the first step in going out and collecting this data once you could do the work under, under the umbrella now of your new company and nonprofit?

9:57I wrote up this really, this like from scratch, a survey platform basically. Wow. And, um, and I just like, I was like, this is amazing. I've never had this much money to collect data. And just started a venture capital. Yeah. And it's just incredible. Like you can, with all the freedom of just having my own company, I was able to get data. Like that was just amazing. And so we started training models on it. Just data from around the world, people undergoing different tasks. Yeah, we had obviously like a lot of overhead before we could actually launch the survey, just like privacy counsel stuff.

10:40and all the compliance stuff. And we actually got all the studies IRB approved because we want to publish, we believe in it, we've published on some of the data. And it was just awesome. So you start Hume, you collect this data, and it all enables you to build this product and really pursue this mission. Tell us what you're doing, tell us what the product is. And of course, we want to get to some of the recent announcements, but would love to hear what you sort of initially set out to do and offer as a product. Hume's mission is to be able to take large-scale data and use it to extract information about what determines whether people are happy or sad, and then use that to optimize AI models for what makes people happy over short and long periods of time.

11:30And you can only do that if you have like measures of the objective proxies of human emotional experience because it's scalable. Otherwise, you'd rely on human labels. With reinforcement learning for human feedback, it's a version of that that just relies on human labels, but it's not as robust. It doesn't involve people in their actual circumstances. It's human raters' opinions, and those raters are just like an arbitrary group of people, not necessarily experts at what people are asking about and the raiders will generally say like oh like this is a very carefully thought out response they're biased toward like these more verbose inoffensive responses that that do a lot of um of couching and like different uh and concert like non-controversial ways and it's just like it makes them boring honestly because at the end of the day these these models are not trained for you to have a good experience they're trained for these raiders to give a positive rating there's a there's a difference there and the model ends up being more sycophantic too.

12:32There's a lot of issues with it. What you want is in the application, when you generate a response, for that response to be optimized for your happiness as a user, like that's what the model should be optimized for. Right. And you're saying there's only so much it can know about me for me typing in a text box. Yeah. Yeah. I mean, even if they have that data, the text is a very impoverished modality. It's very narrow and you're not usually going to tell the model this was a bad response. You're just going to keep going, right? When you're having a conversation with somebody, it's just constant feedback.

13:12There's like every single word has prosody in it. And then you're looking at them, they're back channeling. You're also seeing their facial expression. And it's just constant feedback on the quality of what you're saying. We don't need that when we're texting other humans because we're human and we also have a lot of intuitions that come from our experiences about how that text is going to be received by the other person. Language models don't even start with that, right? So not only is it that the interface is impoverished, but they also don't learn from the kind of data that humans have, first of all, by being human and being able to simulate each other's experiences, but also by living with each other and learning from people's reactions.

13:53It doesn't really understand what makes people happy or sad. Why is it so important for the LLM, for ChatGBT to know if I'm happy or sad when I'm using it? It's important for it to be able to predict whether its response will make you happy or sad. So in order to solve that prediction problem, it has to have in its data set, lots of data on whether responses make people happy or sad. But without that, like as an interface, that is its core sort of goal. And in many ways, you could think of this kind of emotional intelligence as the central capability that interfaces need, which is you give it a query and it's trying to come up with a response that makes you happy.

14:35And if you ask for something, the best response is going to be the thing that you asked for, your intent. A lot of times there's a ton of ambiguity around what your intention actually is. In most cases, that's true. And most of the data sets they train on, actually, that's not true. They train on data sets where the intent is very clear. Or if it's ambiguous, the rater has to determine whether this was your intent. But in reality, in an application, you have a very thin slice of behavior used to indicate what you want. it's not always you don't always know what you want even like if you open up Facebook you don't know what you're looking for it's just the news feed or if you open up Twitter or anything right and even if you do ask for something you want that explanation to be like it's implicit in everything you ask for that you want the explanation to be interesting and not boring right like that's just an implicit part of it that the language model doesn't understand at all or you want it to be funny or amusing or entertaining in different ways.

15:37So, so empathic AI, like an empathic AI could in theory, just be far more effective at making you happy or sad in its response, because it has all these other signals that it can understand and synthesize. Can, can a language model, a pure language model in a text box be augmented by a products like Hume to have this capability? Or is the modality of text completely broken for this type of interaction? So there's the stage of understanding expression, which won't affect the language it uses in response to text. But then when we're doing reward modeling, meaning deciding how the model should form its responses, essentially it's where reward modeling is, beyond just having the capabilities of form different responses.

16:24we're using people's reactions from conversational data and we're able to optimize the model to bring about the right reactions so if you want the model to be funny here's a concrete example like you it would it's helpful to know in conversational data in many millions of hours of conversational data what it is that makes people laugh right and that's something that the model can learn from in the kind of data that we train on. So then becomes a better text model in some ways. Kind of what we talked about earlier, where I'm learning through my non text message based interactions, you training these models with the data that you have is sort of like supplementing that same type of interaction trained into the language model.

17:12It's additional reward modeling data. So even if like, like let's say you have a giant call center and you have a lot of customer service calls and you're a company that provides cable internet or something, something mundane. And a lot of the customer service calls revolve around them walking through different steps to address your issues. And there's like one step in the script that just always frustrates people. You would not know that necessarily from the language alone, but if you have this other data set, you can actually update the model. So it doesn't do that anymore, either in voice interface or in the chat, if people also have a chat interface to that same model.

17:51So it's better in what might appear to be a mundane way. The next logical conclusion then is, okay, so you can take this data, you can take Hume's data, Hume's product, and you can use it to improve a language model by feeding it this data, training it on this data. But then could you take it even a step further by changing the interaction model completely to voice or to video or maybe even video. Totally. So the other side of it is there's so much in our query, a voice query, that language models miss. And part of that is simple things like word emphasis and like when are you done speaking so that it can actually respond in time.

18:32And if you're confused, like what are you confused about? Are you confused? Are you bored? Are you frustrated? Should the model be apologetic? Should it respond with a better explanation? like what should the model actually respond with so it can do all those things better and then it can formulate the right tone of voice um because here's your tone of voice saying if you're frustrated it can be apologetic and um if you're excited it can match that and if you're bored it can you know like it does these very subtle things that you expect a model to do or you don't expect a model do, they expect a human to do that, that naturally form part of your expectations for how anything you're talking to should respond.

19:15And that's just lacking in other models. It's as if like the voice doesn't understand the language. Obviously, like the number one consumer product in the world right now for AI is, is ChatGPT. And it's all happening through a text box. That would lead me to believe that there's like a future world at the application layer that is potentially far more interesting or far more suited, I would say, to AI than the products that we're used to interacting with today. Do you agree with that? And I guess if so, how do you see the application layer evolving maybe beyond the text box? Yeah, I think if you want to put AI into your application, voice is just a better way of doing it.

19:55So it won't just be like one app. I think most apps that use AI will be using the voice. And there's a few reasons for that. One is that voice is faster than typing, like five times faster almost. The other is that while you're talking to something, you can also be looking at other things. You don't have to be like tracking text and doesn't have to take up any space in the interface. So if you have a product that has its own interface, which most products do, the voice is a much more convenient modality for interacting with AI in that product than having to add text in, especially with long text responses.

20:30In addition to all of that, it's personalized, it's friendly, it's something that you want to talk to. Yeah. Yeah. And I think that last thing you said, it's something you want to talk to in this future interface. Definitely true for the demo you launched. Do you call it EV or EVI? What is it? We call it Evie. There was a vote and I voted for Evie, but most people voted for Evie. So we're going to Evie. Tell us about Evie. So Evie is the first way people can experience our product. It's actually an API. The demo is like the first public demo of Evie. And it's a API that combines transcription, language modeling and text generation through our models that we've trained and is able to modulate its tone of voice based on your tone of voice, do all the things that we want it to do.

21:16Modulate what it's saying based on your tone of voice, modulate its tone of voice based on what you're saying. So it's like fully cross modal. The effect of that is that it feels like it understands you. I think from a user experience perspective, it's different. And so we see people having long chats with us, which we did not expect. Like the average is nine or 10 minutes. And then people are having like 40 plus minute conversations with it and actually writing in and being like this made my day. Like this was what I needed. Yeah. How do you imagine companies are going to leverage this? So is it a layer that they will add over their own models or they'll take this plus, I think, the model that you offer as an interface for their application and what will those applications be?

22:01So any app that uses an LLM can use pretty easily our interface with the same configuration. So we have our own language model, which generates the actual speech, but we can call APIs to get the reasoning capabilities of the most frontier models. We can integrate web search very easily. And then as we output the response, we can also output tool use. So we can navigate a web page, we can navigate an app, we can call any API that's called by the front end of the app. So this could be integrated pretty easily into any application. We actually have it like as a widget on our website you can talk to and it navigates.

22:42That's awesome. So like what are the types of companies or products you think are gonna be the first to adopt this? Anyone who really cares about the user experience and wants to integrate AI into their product. And I think there's quite a few of those. Where user experience makes a huge obvious difference, especially is like health and wellness. Like if you have a patient, they want to talk to something that's more comforting, that understands them better, not just a patient, but anyone with a health and wellness app. So there's a lot of signups for like various kinds of health and wellness apps.

23:18There's also like other times when you want to be comforted by what you're talking to, like financial services. Like you want to be talking to something that you trust and it has to modulate its tone to kind of meet you where you're at because there's a lot at stake. And so a lot of these, even banks are launching financial assistance and the user experience matters a lot. And there's high stakes. Got it. So financial experience app, maybe instead of chatting with a text-based client, now I'm voice interacting with an AI agent that's maybe, I don't know, helping me troubleshoot or maybe making recommendations for me.

23:57and it's talking to me in this really empathetic voice like a human would. Yeah. Imagine you're like it's tax season. So imagine you're doing your taxes and you're confused by some things. You're frustrated by other things. And there's a huge interface with like tons of forms to go through. Imagine if you could just like say, hey, take me to where like this is decided and it takes you to the right page and starts filling out the form for you based on your data and it asks you in real time, like, is this right? Is this right? Does this look right? What about robotics? Like, I feel like, you know, there's obviously a lot of talk about robotics right now.

24:33It feels like this could be a huge piece of that puzzle, right? Where you have robotics that are now becoming more human-like through this type of interaction. I think the voice is extremely important for obvious reasons for robots and wearables, where like, or devices where you really want to be able to look at things and talk at the same time. like you don't want to have to take your eyes off of what you're doing and read so that's a pretty obvious use case for voice and user experience is incredibly important being able to be understood quickly and feel like it's understanding you and you you know it understands you because it's reflecting it in its patterns of voice and and and just the ability to transmit information about what's going on in the world, being able to react to things.

25:25And it knows because it sees the same things that you're reacting to and what you're looking at, like what you're frustrated by, what's confusing and kind of wordlessly or in very concise way addressing those things, I think is core to the experience. Yeah, you mentioned that you can talk a lot faster than you can type. Obviously, there's a lot of talk of AI being sort of this next platform shift. but as we talked about a few minutes ago, like we're still typing into these text boxes with, which feels like inefficient means of computing, you know, compared to the GUI as an example. I wonder if, if, if voice is maybe the real unlock for AI becoming the next platform shift and, and does it require something like this to get me to want to interact with it in a human like way?

26:16Do you agree with that? And how do you see computing evolving based on this? On the desktop or on a mobile device or maybe something else, some other type of hardware or form factor? People are going to move to more spatial computing where it's with you all the time. It knows you. It learns from your... So learning from your patterns of speech is important for that because it becomes more personalized. Talks in the way that you want it to talk. And it's present with you in different places. So voice has to be there for that. Like you don't want to have to pull out text. I think that this modality really requires the end to end understanding of tone of voice in order to make the experience something you actually want.

27:03It just doesn't feel good to talk to an AI for extended durations of time that sounds uncanny, that doesn't sound like it understands your voice, that doesn't sound like it understands what it's saying. And naturally, we kind of expect, we build into our expectations for how somebody's going to respond. We just build in this expectation that they're going to understand our tone of voice. So when that's not there, it just seems weird. Yeah. And it feels like that's been a big limitation of Siri for many years, right? I mean, Siri, it's almost like a running joke that it's not super performant, but also you You have to kind of talk to it in a very specific way to make sure that it understands you.

27:44I feel like Alexa and Google Home got a little bit better, but all of these things still don't, I think, fully grasp maybe the intent behind what you're saying because of the reasons you mentioned. Like, it doesn't really understand tone of voice. And so I could see like how this would solve that. I guess the other half of that, which I think also leads to something you're doing, is this then notion of being able to actually act on behalf of me as an agent, right? should be able to talk to Siri or Alexa or Google Home or any application, like you're saying, via AI and have it take an action. And that's something that you alluded to that you're now doing as well.

28:20Is that right? Yeah, so we don't really build much of the tool use ourselves. And I think developers are already building that on language models. So it's just a matter of being able to unlock that capability. So if you send your configuration, we can call any API that you're calling and we can give function calls that the client executes that you need. Anything that's happening client-side, we can tell Eevee that this is happening and it can call that API at the appropriate moment. So that's going to be a really key thing to enable developers to build. We're not going through people's apps and looking at all their API calls, but the process of doing that is not that difficult.

Read the full transcript

29:03So developers can do that just by prompting our model and giving it the right function calls that it can use. I think that's generally going to be the way things are done. Now, until recently, I don't think language models were able to carry out tool use very reliably. I think that's changing really fast. Why is that? Why couldn't it? Just the amount of training required to get it to be able to do it well. And you also want it to be pretty reliable. Sometimes 90 % is not reliable enough. For us, we're building certain tools into the backend that are very reliable, that people might want frequently, like web search.

29:37But for the most part, we're leaving that to developers to do. And the frontier language models can call tools at the right times. So we can actually hook into frontier language models when we need to. We also build on pretty good models ourselves. So they're also capable. But if we want to use the latest large language model trained by OpenAI or Anthropic, we can route the request through that. And then they can handle the function calling, which then goes into the context of R. Do you want to be able to route any request from within Hume? Or are you saying it's really up to the application developer that leverages the Hume API to figure out how to route it and to plug in the API?

30:23It's up to the application developer and I think that that's going to be true because applications all have their own requirements and especially what you're doing at the front end is going to be unique to your application. So the ideal is, I mean like what you could do is, you could put your entire, this is extreme and you wouldn't actually need to do this, but you could put your entire code base into the context window of say like the latest frontier language model and it would be able to look up all the different API calls you're making. Maybe you want more control than that, but it could potentially just right off the bat be able to do those things.

31:03In reality, you don't need your entire code base in there. You just need to know what the configurations required for the various function calls your application is making are and then just tell the LLM, and this is the beauty of LLM, so you can just tell them in natural language when it's appropriate to call that function. Two-part question and maybe a bit philosophical. How do you define AGI and do you feel like empathic AI is an important ingredient to reaching AGI? I mean, I would define AGI as AI that, I mean, the functional definition is AI that can do basically any task that a human can do equally well or better, right?

31:48I think that we'll get there before. But if we get to a point where AI can do those things, but it just doesn't have any sense of concern for our well-being, it's not going to look good for us, I don't think. And the reason for that is that you'll have agents that are out in the world that you tell, hey, decide on how to raise the value of the money in my bank account to the greatest extent possible in the next week. And that's a pretty ambiguous request because it could just do that by taking out a lot of debt, like really bad terms for you. But it has to have a background understanding of your preferences and alignment with your wellbeing to be able to carry out most requests in the way that you actually want it to.

32:36So I think that that's where emotional intelligence comes in and optimizing for human wellbeing really comes in. Like if you train and model on enough data with indicators of people's emotions and their experiences and whether they're doing well, mental health and so forth. If you train on enough data, the model is going to have a much better understanding of four steps into the future, how its actions affect your well-being. And that's what you sort of want it to have. And that becomes kind of a background for everything it does. That's a background for how humans interact with each other, too.

33:11Most of our requests are fundamentally ambiguous to each other, but we rely on the fact that the person on the receiving end is a human who can simulate what would happen to you if they carried out your request in various ways and figure out, okay, this is the best way to do it. And so we call that common sense, but it's actually part of emotional intelligence. And if AGI lacks that, I don't think it's going to be, I wouldn't say that it's like, I'm not a doomer. I don't think it's going to take the world down because I think we just won't use it if it's not, if it's not working the way we want it to.

33:44It just won't work the way we want it to. I had another conversation on this podcast with Gustav Soderstrom, who's one of the presidents of Spotify. And I asked him about AGI and his answer was, was actually interestingly, very similar to yours and that he felt like with enough data, AI would learn from our empathy and therefore it would be empathetic towards human beings. But now based on everything you're saying, this entire conversation, really, it feels like a way for it to have that data is actually through a data set like Hume's, right? It's not just going to get that by looking at every written word that's ever been published on the internet.

34:23that it's gonna need to understand our emotional behaviors, our facial expressions, how we feel. And in that sense, it feels like Hume could be a very, very important piece of that puzzle. Yeah, we hope so. I mean, that was sort of the motivation from the beginning. And yeah, you could get some of it from just language, but it's not going to be the most capable capability that it has from just language. When you layer in other modalities, a lot of what you're bringing in is relevant to emotional experience. And so you're layering in that data with every word, and it becomes a much bigger focus for the model, and it becomes one of its greatest capabilities.

35:08It's his ability to predict what makes people laugh, what makes people cry, what makes people confused or satisfied or awe-inspired. It's able to predict all of that because it can predict your reaction to something in the next minute or hour or day. That's one of the main things that is added. And when you just have language, I think 90 % of that is lost. We talked a little bit earlier about hardware. And I think a lot of people are talking about hardware right now. We're starting to see all these new devices come out. humane, the rabbit, obviously, I'm sure Apple is gonna get involved here. And it feels like voice is a huge, will be a huge part of this new hardware frontier for AI.

35:52Like what are your predictions for hardware and AI? Yeah, I think people will want hardware, and I don't have a strong opinion on form factor yet. I think that there's a lot of cool things that might be the right form factor. I think a necklace could be really cool because it can listen to your voice and it has a pretty strong signal of that no matter where you are, even in a noisy environment. The iPhone could be fine. Like just like having it built into your phone. Or the watch. Or watch. Yeah. Various wearables could all be good form factors for that. I think the important thing will be for the AI to understand your voice and also have a sensory awareness of what's going on in your context.

36:37So that might be made more powerful by having a front-facing camera, which would constrain the form factor a lot. I'm not sure. But definitely just being able to listen to what's going on is a huge part of it. What do you think about, I mean, I'm not the first person that's asked this, but what do you think about the privacy implications of that, of now these devices listening in a way that we haven't had before, but not just listening, being able to then feed a new type of data set and model like the one we're talking about that hasn't really existed before. How are humans going to kind of get over that hump?

37:15Yeah, I think that that's going to be a real concern and they have to trust the AI that's processing their data. So there's kind of two layers to the question. One is sort of what is the AI doing with your data? And the other is, is this data protected from like being exposed to other humans, right? And so those are kind of two different questions. And if you've solved the second one, which is privacy, which I think you already have tons of data on your phone, you already have tons of data on your computer, you're already kind of assuming in order to live in this world that your privacy is protected by the devices that you're using.

37:52So that's sort of an assumption. So once you've solved that part, the other side of the question is, is the AI doing something with your data that you want it to do? And so you really need to trust these models, these algorithms. How do they build that trust? I think that in order to build that trust, they have to demonstrate that they make decisions on the basis of what's right for you and not some third party. And so we have to move away from optimizing algorithms for engagement or for just selling you things or any third party objective that could conflict with your own well-being. I think models that do that will not be the models that you want to deal with in everyday life because they're super intelligent.

38:34And so you're negotiating with something. You have so much leverage over you, right? And you actually want something that's on your side to mediate that interaction. So I think that a lot of this will depend on having an AI that you trust to be optimized for the right things using your data. What's next for Hume? There's a lot of improvements coming to just what we have. So that's the first thing is that people have enjoyed using the demo. I think that what we're launching in the next month or two, first of all, we're launching the API, which is another big thing. So developers can start using that.

39:06And then we've had a lot of requests for that. But even just the experience of what we have so far is going to greatly improve and people are going to feel the difference. I think we do want to make it available to people in a kind of a format that end users can use directly, not competing with developers who are building applications. All the function calling, all the tool use, we leave that to developers, and I think that that's going to be part of the future, is developers building their own tool use capabilities into these models, and it's not just one centralized company doing all of it. But we do want to be an interface that people trust to access those tools and that people trust is on their side that can optimize for people's well-being within the application that developers can fine-tune for people's well-being within their applications, which I think will build trust over time.

40:01Maybe the last thing is, how do you assemble a team for a company like this? It feels like you would need so many different disciplines and a wide range of experiences. What does this team look like? And also, are you hiring? And where can people learn more about that? We are hiring. We have a job board and we're looking for a lot of different profiles right now, but particularly people in ML training and inference optimization and distributed training and also many other areas. So even if you're new to ML, there's lots of roles. The team is, I think it's a pretty interesting team. We have PhDs in computer science, but we also have PhDs in psychology and neuroscience.

40:55And we have a lot of people who have recently entered into the AI space and just show a lot of enthusiasm and talent. And I think people joining the team end up learning a lot from a lot of different perspectives. and that's really fun to see. Alan, this has been a fascinating conversation. I feel like I learned so much and I hope the audience has as well. Thank you so much for the time and I hope we get to do it again sometime. Of course, thanks for having me on.

41:30Thank you so much for listening to Generative Now. If you liked what you heard, please do us a favor and rate and review the podcast on Spotify and Apple Podcasts. That really does help. And if you want to learn more, follow Lightspeed at Lightspeed VP on YouTube, Twitter, LinkedIn, Instagram, or anywhere else. Generative Now is produced by Lightspeed in partnership with Pod People. I am Michael Magnano. We will be back next week with another conversation. Thanks so much.

From the publisher

What if AI could understand what we, humans, are feeling? This week on Generative Now, Lightspeed Partner and host Michael Mignano talks to Alan Cowen, CEO & Chief Scientist of Hume AI. Hume creates empathic AI that learns our preferences from our vocal and facial expressions. Their goal is to maximize our happiness and quality of life. Hume is now announcing their API for EVI, Hume’s Empathic Voice Interface. 

The conversation covers Cowen's journey from being a researcher to founding Hume AI, the importance of emotional intelligence in AI for quality human interaction, and the potential to transform user experiences across various apps and devices. Plus, we ask how empathic AI could impact the road to AGI.


Episode Chapters
(00:00) Introduction to Alan Cowen & Hume Demo 

(01:38) The Genesis of Hume AI: From Research to Startup

(04:01) Affective Computing and Its Impact

(10:55) Hume AI: Bridging Human Emotions and Technology

(15:37) The Future of AI: Beyond Text to Empathic Interactions

(20:37) Introducing EVI: Empathic Voice Interface

(21:46) Real-World Applications of Empathic AI 

(31:19) The Potential Role of Empathic AI in Achieving AGI

(36:53) Trust and Privacy

(40:02) Opportunities with Hume AI

(41:18) Closing Thoughts 

Stay in touch:


The content here does not constitute tax, legal, business or investment advice or an offer to provide such advice, should not be construed as advocating the purchase or sale of any security or investment or a recommendation of any company, and is not an offer, or solicitation of an offer, for the purchase or sale of any security or investment product. For more details please see lsvp.com/legal.

More from Generative Now | AI Builders on Creating the Future

All 90 episodes
Alan Cowen: Creating Empathic AI with HumeGenerative Now | AI Builders on Creating the Future · 42 min
Listen in VO