In short
Big Technology Podcast: Episode Summary
Episode Title
Teaching AI To Read Our Emotions — With Alan Cowen
Host
Alex Kantrowitz
Guest
Alan Cowen, CEO and Founder of Hume AI ---
Episode Overview In this episode of the Big Technology Podcast, Alex Kantrowitz interviews Alan Cowen, who discusses the significance of emotional intelligence in AI systems. Cowen's company, Hume AI, is at the forefront of integrating emotional understanding into AI by analyzing voice and facial expressions. The conversation covers the theory behind reading emotions, its practical applications, and the potential future impacts on various sectors such as customer service and even understanding animal emotions.
---
Key Topics Discussed
- The Importance of Emotional Intelligence in AI
- Understanding Human Emotions:
- Emotional intelligence is essential for AI to respond appropriately to human needs beyond the literal interpretation of text.
- AI's ability to gauge emotional reactions can enhance user satisfaction and overall interaction.
- Communication Modalities:
- Emotional understanding encompasses various forms of communication, including voice, facial expressions, and text.
- Challenges of Emotion Recognition
- Text Limitations:
- Text communication is often ambiguous and lacks the emotional depth conveyed through voice or visual cues.
- Punctuation, capitalization, and context are insufficient for fully understanding emotions in text.
- Innovative Approaches in Hume AI
- AI Bot 'Evie':
- Evie can recognize and respond to emotional cues in voice, marking a significant advancement over traditional text-based AI systems.
- The model incorporates over 756 dimensions of vocal modulation and facial expressions to enhance understanding.
- Historical Context of Emotion Studies
- Evolution of Emotion Research:
- The conversation traces the lineage of emotion study from Darwin to contemporary researchers like Paul Ekman.
- Emphasized the shift from viewing emotions as discrete categories to recognizing a continuum of emotional expressions.
- Applications of Emotionally Intelligent AI
- Customer Service:
- AI can streamline customer interactions by integrating emotional understanding, potentially replacing traditional human agents.
- AI could enhance user experience by resolving issues within applications directly rather than through conventional call centers.
- Understanding Animals:
- Cowen mentions potential studies on using AI to interpret animal emotions, especially in social species like primates.
- Future of AI and Emotional Intelligence
- Integration into Daily Life:
- Cowen envisions AI systems that can learn preferences and provide tailored responses, improving personal engagement.
- The evolution of AI could lead to a more symbiotic relationship between humans and machines, emphasizing emotional connectivity.
- Ethical Considerations
- Care in AI Design:
- The design of emotional AI must prioritize user wellbeing, preventing manipulative or intrusive interactions.
- The conversation highlights the necessity of maintaining ethical boundaries as AI capabilities expand.
---
Key Takeaways
- Emotional AI is the Future: As AI continues to evolve, the integration of emotional intelligence will be crucial for developing more satisfying user experiences.
- Broader Applications Beyond Text: Understanding emotions through voice and facial expressions can lead to significant advancements in various fields, including healthcare, customer service, and animal communication.
- Ethics and Responsibility: While the potential benefits are vast, developers must approach emotional AI with caution, considering the ethical implications of creating emotionally responsive systems.
---
Conclusion This episode of the Big Technology Podcast provides a comprehensive look at the intersection of AI and emotional intelligence, featuring insights from Alan Cowen on the transformative potential of AI systems capable of understanding and responding to human emotions. As the technology progresses, the implications for various industries and human interaction are profound, marking a significant step toward more intuitive and responsive AI solutions.
---
Listen to the Episode For those interested in exploring the full conversation, listen to the episode on [Big Technology Podcast](https://www.bigtechnologypodcast.com).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Let's talk with the CEO of the leading AI startup for emotional intelligence, looking at where this exciting new discipline of AI is heading right after this.
0:12you're used to hearing my voice on the world bringing you interviews from around the globe and you hear me reporting environment and climate news i'm carolyn beeler and i'm marco worman we're now with you hosting the world together more global journalism with a fresh new sound listen to the world on your local public radio station and wherever you find your podcasts
0:40Welcome to Big Technology Podcast, a show for cool-headed, nuanced conversation of the tech world and beyond. We have a great show today going into one of the more interesting and emerging areas of AI research and application. I think it's going to be fascinating. We're probably going to see this stuff packed into basically every AI that we touch going forward. And we're going to speak with somebody who's really at the cutting edge of where this is all going. Alan Cowan is here. He is the CEO and founder of Yume AI, which is a company with a bot that we've talked about on the show here in the past, one that can sense the emotion in your voice and reply in kind.
1:15It's very fascinating. And I cannot tell you how excited I am to have Alan on the show. Alan, welcome to Big Technology. Hi, Alex. Thanks for having me on the show. We're going to speak with Dwarakish Patel in the next couple of weeks, but he had a podcast episode with Mark Zuckerberg basically right around the introduction of Llama 3. And they were talking about the cutting edge of AI research. And Zuckerberg says emotional AI for some reason is something that I don't think a lot of people are paying attention to, but there's a lot of potential there. One modality that I'm pretty focused on that I haven't seen as many other people in the industry focus on this is sort of like emotional understanding.
1:55Like, I mean, so much of the human brain is just dedicated to understanding people and kind of like understanding your expressions and emotions. And that's like its own whole modality. What do you think he was getting at when he said that? I'm not entirely sure what's in his mind, but it's definitely the right problem. I think the reason it's the right problem is that when you're trying to address people's, let's say, concerns with AI, and also when you're trying to just meet their requests, the most important thing is emotional intelligence. It's the ability to take people's preferences into account in your response.
2:32And so emotional AI is, first of all, all AI is emotional in the sense that like all AI has to do this. But AI that's specifically trained to do this is going to learn to produce responses that make people happier. And that's the fundamental objective for AI. And why is that important in terms of advancing AI? So if you're talking to ChatGPT, you're implicitly saying, okay, I'm going to give you a way to make me happy, which is to fulfill this request. Oftentimes, you don't even have an explicit request. And when you do, it's actually fundamentally ambiguous, right? So the goal of an interface is to take this ambiguous slice of behavior.
3:13And it could be language or it could be language plus audio plus facial expression plus eye contact, whatever behavior it has access to. And it's trying to take contingent on that the best estimate of what might make you happy if it responds in a given way. So that's literally what it's trying to do. It should be able to learn that by understanding human emotional reactions to things. That's the fundamental way that humans learn this stuff. Like we can either simulate what your emotional response will be to me doing something for you based on putting myself in your shoes, or we can look at your emotional reaction based on that, figure out whether I did the right thing or not and adjust.
4:00And those are really the ways that we learn to make each other happy. So much of that is actually done in voice, done in expression. And in text, it's so much harder. Let's start with the hardest problem and then kind of work back from there. How do you assess out emotion in like a text exchange? In text, there's some cues, right? And most of the time, we're not actually sussing out emotion in texts. We just have certain conventions that we use those conventions to dictate how we communicate. And so we know that this is a very narrow channel through which we can express our desires. And so we know that there's fundamental ambiguity.
4:37So we just basically limit what we try to do. We don't try to do too much with text. We don't try to bite off more than we can chew. What are some of the conventions that you'd say we rely on with text? Punctuation and capitalization and where you put periods and when you want to add an extra Y to the hay or like all these different conventions. And they just have like these kind of stereotyped meanings. So it's a message passing system. It's very frustrating to correspond by text if what you actually are trying to say is anything more complex than that. What's coming across here is that it's actually remarkable how limited this current set of large language models are, given that it's all text and it's still blowing people's minds.
5:24But there's really a limited range in which it can be useful. Is that right? Yeah, and it's never seen people's emotional reactions in any other modality. It learns some things from text about what people want and don't want, but it actually doesn't have that layer of understanding of how is this going to make people feel. So even via text, it seems to be missing something. And yeah, it can solve problems. If you specify the problem in exactly the right way, for example, in the way that these LLMs are evaluated, you just literally ask it a multiple choice question. If you specify it with that level of constraint, then they're really good.
6:00But anytime you want to do something more open-ended, it's a relatively unsatisfying experience. So this is why people are gauging how good these things are by asking them to take the LSAT. It's like not an accident. It's like this is sort of the only way to evaluate them within their constraints. And that's more or less how we evaluate. So like the biggest eval is MMLU that people use. And it's literally you can get each question right or wrong. And then you just look at the percentage they got right. It's basically a big multiple choice test with some other kinds of questions in it. When I first started using your product, you have a bot.
6:33It's called Evie, E-V-I. And when I first started using it, it's kind of in demo mode now. And the API just opened. We can talk about that. But I was like, oh, this seems like a large language model in some ways. And that like it's giving some of those kind of wink wink, like as an AI assistant, this is my feelings. As an AI system, my purpose is not to replicate the full human experience, but to engage in meaningful and empathetic conversations. But it's voice. And it's now coming into focus that you decided voice because there's so much more signal and also so much more realm of possibility in terms of interacting with this than you can with a text bot.
7:19there's so much more data that comes in via voice and can go out via voice, then we can really take a next step in terms of the way that generative AI is moving. Previous technology relied exclusively on transcribing the voice and taking away all that extra meaning that's in your vocal modulations. So this technology didn't exist before Hume. When you add that in, what's remarkable is that the model can actually learn what vocal modulations mean on its own. You just need to add it in the right way. We have these expression measurement tools that we've trained that are extremely accurate. They can actually encode 756 dimensions of embedding of vocal modulation and 756 dimensions of facial expression per word that you're saying.
8:03So there's these incredibly rich embeddings that we're adding in that add a lot more information. And the model can actually learn that. And it turns out that it's just more satisfying to use when it has that data. Take us through very briefly a sort of ride through the study of emotion in humans, where it was lacking and how it's now being applied into robots. Or maybe not that briefly. I feel like that's a tall task to do in a minute. I mean, it starts with Charles Darwin. I won't go back that far. But basically, he was sketching out facial expressions in both man and animals. And he has a book called The Expression of Emotion in Man and Animals.
8:41And he recognizes the complexity of it and the nuance. but somehow that kind of was lost over time nobody really talked about that book for a while in the mid 20th century there was Paul Ekman who decided he would tackle the question of what is it about facial expression that may or may not be universal and so he cataloged these six facial expressions as an arbitrary test more or less of what these broad expression stereotypes might mean to different people. And he chose them because they're very different. He only chose one positive expression. He chose five negative expressions. Wait, which ones are they?
9:20Anger, disgust, fear, happiness, sadness, and surprise. Wow. Surprise could be good, but it's definitely a largely over-indexed on bad feelings rubric. It was bad in his conceptualization of it. Yeah. So positive surprise is good. negative surprise is bad. And then there's also related epistemic emotions like awe and interest, which are good, or confusion, which is bad. So to varying extents with varying levels of arousal and excitement. Anyway, but none of these distinctions are in those images. And for some reason, people just stuck to those six categories for a while. And there was this countervailing narrative that actually there are fewer granularities in emotion that people recognize, that actually it's just valence, like positive, negative versus arousal, calm or excited.
10:10I don't know why that was the countervailing argument, except that maybe people were just limited by the amount of data that they could collect. And this was like the closest thing to data driven that they had, which is like, if you do dimensionality reduction on a hundred samples, you get positive, negative, uh, low arousal, high arousal as your dimensions. And wait, arousal in terms of what is arousal mean? Arousal means like excitement, not like physiological arousal, meaning like activation of the autonomic nervous system. Okay. Good to clarify that. Okay. Yeah. And that's how, I mean, there's lots of dimensions to that.
10:46Obviously, like one of them is actually sexual arousal. Like that would be different. Like we recognize that it's very different from the arousal that goes with fear and like running for your life. Right. And yet in this theory, they're compressed. I know. Well, it depends who you ask, but okay. Continue. you can kill that one can cause the other i don't know well anyway that was it's true this the the not to go too far in the digression but there is scientific literature that says danger and arousal are connected i mean esther pearl had a whole book on it called mating in captivity that's excellent but yeah well yeah it might not be danger but not our specialty novelty right yeah and there was there's misattribution of arousal which is this old experiment where you take somebody on a bridge and is very exciting.
11:36And then you test whether they basically are aroused. And yeah, they are. Anyway, or whether they're attracted to you. And this is not something that necessarily replicates in every condition. OK, so there's lots of different kinds of emotion. They get reduced down to these six basic categories or two dimensions for a very long time in research. And a lot of that has to do with the fact that you can only collect so much data. The data analysis tools that were available to people for a long time were very coarse. For example, there weren't even really computers that psychologists were using to analyze data.
12:14Incidentally, psychologists did, like the field of stats sort of came out of psychology, but that was earlier. And then, you know, when it came to actually applying it to lots of data, that was much harder and that didn't come until much later. So now data science comes around sort of in the 2000s and becomes a really big thing. And psychology was one of the last places that it was applied. So it starts to get applied to neuroscience and biology and genetics and all these other areas. But people didn't really have the tools to analyze human behavior until relatively recently. and part of that is like even when you have the the data analysis techniques the data collection techniques are difficult so what i pioneered in my phd in psychology and while i was working for google and facebook was a new kind of data collection that you could then use to collect enough psychologically controlled data to apply data science to psychology data and that's that's ultimately what gave us a lot of surprising results about the dimensionality of emotion.
13:29What were the surprising results? What you can see clearly is that people make really granular distinctions between lots and lots of different expressions pretty reliably. So there's differences between an expression of awe versus surprise versus fear versus confusion, interest. All of these expressions are reliably recognized and so we just started to map them out. We realized they're not discrete categories, that they're actually continuous and can be blended together. We realized that the number of dimensions that it takes to represent these things is large. It's not like you can reduce it down to valence and arousal.
14:09That actually captures like 20 % of the variability in people's ratings that's consistent across different raters. you actually need you know in facial expression over 30 dimensions and speech prosody the tune rhythm and timbre of speech you need over 18 different dimensions to represent how people are able to conceptualize speech prosody and in vocal verse like laughs and sighs and screams you need at least 24 different dimensions and this is just what people explicitly recognize it actually goes beyond that when you look at what's implicit in people's responses to things that is not well-verbalized, or what's implicit in people's conversational signals that's not verbalized explicitly.
14:52So there's a lot of different dimensions that had never been classified before. So you've never classified these dimensions of behavior or been able to measure them. There could never have been a science of what they mean. And so this sort of opens up the door to actually understanding human expressive behavior in the natural world, in the real world and understanding what it means. And we got a ton of publications out of that pretty quickly. And then you also seem to have been able to study how much small variations in these expressions or even tone of voice can tell you. And talk a little bit about that, especially the tune and the timbre of the voice and what we can learn from each of those.
15:34We created new machine learning models that can reliably distinguish among all these different dimensions of vocal modulation that people form and make it distinct from the language, the phonetics of what they're saying. But that was a challenge before. So we solved that. And when you can do this, you start to see a lot of meaning in people's voice modulations that is just completely not present in the phonetics or present to a much lesser extent. Like, are you, and it can be simple things like, are you done speaking? And obviously, like there's a lot of emotional dimensions that get added to a lot of words or every word is every word carries not just the phonetics but also a ton of detail in its like tune rhythm and tamper that is very informative in a lot of different ways you can predict a lot of things you can predict whether somebody has depression or parkinson's to some extent really perfectly yeah so like mental health conditions can be predicted to some extent you can predict in a customer service call, whether somebody's having a good or bad call, much more accurately by incorporating expression than just with language alone.
16:41So pretty much any outcome, like it benefits to include measures of voice modulation and not just language. And what is phonetics? Phonetics meaning like the underlying words, but specifically when we form words, there's a phonetic representation that gets converted to text or that's converted to semantics, basically. And that's what the transcription models are doing, is that when you convert something to text, you're actually only relying on the phonetic information, meaning the stuff that conveys word forms, and you're throwing out everything else. You're throwing out stuff that might convey, for example, emotionality or tone of voice, et cetera.
17:31It's so interesting because I'm thinking now about like, so often what I'll do is I'll take the transcripts of podcasts and dump them into Claude and start talking to Claude about it. But of course, it's only getting the words and not the emotions in there. So there might be parts where I thought, or it might be interviews which I thought were like particularly like rough or really exciting but it can't fully pick that out because it's just seeing the text and there is so much meaning contained in the words that are said uh that that you need deeper tech and deeper expertise to pick out yeah i think some people forget like there's a lot of density to the information in audio that like if you transcribe a conversation sometimes it doesn't even make sense to read it.
18:17Like you can read the words and you're like, this actually does not, my brain cannot make sense of this. But then when you hear the audio, like it makes perfect sense to you. Totally. Like even speaking with the media, I remember I was on the first time I was on NPR, I was very excited. I went to the studio, recorded my interview and from the audio, it sounded like a normal conversation. But then I read the transcript and I'm like, man, what were you saying there? None of this makes any sense but it's it is such different meanings when you just look at the text versus like hear like the full express conversation in the meaning so yeah um if you look at political speeches often that's the case i'll say more about that even if the writing's really good you look at like mlk's speeches if you transcribe them and read it it just doesn't have the same effect at all not the same no we're close no emotion to them or less people yeah there's certain people who don't use the right grammar or linguistic forms that would be traditional in writing um either because they're just not good at it or because you know they can depart from that purposely like if you look at like trump's speeches yeah they really when you transcribe them a lot of them are completely incoherent well when you listen to them yeah i mean maybe some of that is true going out on a limb here but but yeah i i do think that it is interesting because when I'll write for spoken word, it'll always be different than the written word.
19:44We are going to talk about how this technology is going to be applied and already is being applied in the business world, but just still a few more theoretical questions for you. Is this type of technology something that we can use to understand animals? Like people talk about using AI to understand whales. Is this something that can be used in that nature? so the same types of machine learning models that we're training can be used for for that our data would not really be that relevant to that right because it's a totally different language yeah and maybe some of the methods the sort of the training approaches that we're using could be helpful for that yeah um but it's a different problem um maybe to you know we have an interest from evolutionary biologists who want to measure the similarity between expressions and different mammals and human expressions.
20:38And so there, maybe there's some, there is some relevance because we can, we can treat it like a human sort of morph it to look like a human face and then analyze the expression and see if it predicts the things that we want to predict. Are you going to do that type of work? We can, but like, it's not really that strong of a model to actually perform inferences. It does show similarities. Like there's dimensions of similarity between humans and chimps. Chimps have an open face model, which means something similar to when humans are laughing and playing, and they use it when they're playing, both physically and kind of in a non-physical touch scenario.
21:18Mice laugh. For real? Yeah, hypersonically. So like you can't ultrasonic, so you can't like hear it at all if you don't turn down the frequency of it so that it's something humans can hear. But then it kind of sounds like a laugh. No way. And you can elicit this by tickling them. If you tickle mice. They laugh? Like this. Yeah. Kind of makes me feel bad that we use mice so often in scientific experiments knowing that they can laugh. Yeah. I mean, they also have a lot of empathy for each other. Like if one mouse is trapped, the other, without any ulterior motive, will hear that mouse screaming and then like go and try to investigate it and get them free.
22:03And if there's a lever, they'll figure it out. Really? Yeah, they're pretty smart, actually. Yeah. Mice and rats are also more related to humans than they are to dogs. Right. Well, that's why we do all this testing on mice. It's because they are close to us. Yeah, exactly. It does make me feel bad. And then, um, do you think that we're going to be able to read like, uh, using technology to be able to understand animals? As someone so close to this, what's your prediction? I think that the, the animals will have some kind of quasi language that they use. And so we know this for, for like, for like primates, for example, that there's different calls that mean like there's a snake below us.
22:49And then if you like play that through a speaker, all they're all look down or like there's an eagle above us and you play that through speaker, they're all look up. Right. And so there's like a quasi language there. I don't think they have the same level of syntax that humans have. And syntax is our ability to form sentences with nouns and verbs and relationships between entities and logic and stuff. And I don't think they have as much of that. Some people speculate that dolphins might have that. which would be pretty surprising i don't know for sure um because they don't have like hands and so the value of that ability like maybe they use it to coordinate like hunting and stuff but it's hard to imagine like why that's important for the dolphin brain to have that capability yeah but you know if they're they do seem to be really smart there's that like scene in blue planet i don't know if you've seen it where the the dolphins and the false killer whales are first they're like fighting like the killer whales are chasing the dolphins and then there's a scene where the dolphins like turn around and they all just stop and they just seem to be squeaking at each other for a while and then after that they hunt together for real i haven't seen yeah it's insane so that is suggestive to me that maybe there is something to the language potentially and ai's ability to decode that it's within our lifetimes we could it's what people are working on it people are working on it it's difficult because you don't really have a grounding for it you kind of just have to like like ideally you'd have really detailed explanations of what's going on to accompany these this language that's being exchanged but like you don't have that that's like that's how we train image models we have captions for the images right If you just had to teach a computer to understand images without captions, it would be difficult to ground that in anything.
24:51Like you'd have a computer that embeds images and can tell whether they're similar. That's fine. But at the end of the day, you want it to be able to take the image and explain it to you. Then you need some grounding for that. You might not need that much grounding if you've done enough compression of the similarities between images first. which is sort of maybe that's what we're hoping to do with dolphins. Fascinating. Okay, I want to talk about the business applications here, including why Facebook and Google would want to employ you to put some of this stuff into work inside their product. So let's do that right after the break.
25:23We'll be back right after this. Did you know your credit card points and miles can lose value to inflation? Credit card companies often reduce the redemption value of your points and miles. Now, imagine a credit card with rewards that can grow in value. With the Gemini credit card, you can earn Bitcoin or one of over 50 other cryptos instantly with no annual fee. Every swipe at the store or gas pump earns you instant rewards deposited straight to your account. Plus, sign up now for a$200 Bitcoin bonus to kickstart your rewards. Visit Gemini.com slash card today. Check out the link in the description for more information on rates.
26:01Again, if you're looking to invest in Bitcoin but don't know where to start, the Gemini credit card makes it easy. The Gemini credit card is issued by WebBank. In order to qualify for the$200 crypto intro bonus, you must spend$3 ,000 in your first 90 days. Some exclusions apply to instant rewards in which rewards are deposited when the transaction posts. This content is not investment advice and trading crypto involves risk. The Gemini credit card cannot be used to make gambling-related purchases.
26:32You're used to hearing my voice on the world bringing you interviews from around the globe. And you hear me reporting environment and climate news. I'm Carolyn Beeler. And I'm Marco Werman. We're now with you hosting The World Together. More global journalism with a fresh new sound. Listen to The World on your local public radio station and wherever you find your podcasts.
26:59and we're back here on big technology podcast with alan cowen the ceo and founder of yoom it is a company you should check out that has this very interesting new bot called evie and also an api that's going to allow companies to take advantage of this ability to determine human emotion through voice and maybe one day through facial expressions. So let's talk briefly about your work at Facebook and Google. So you've done this work. You've basically expanded the number of human emotions that we sort of acknowledge exist and can study and have worked to build that into machine learning models that can recognize them.
27:44What did Facebook and Google want with that type of knowledge and technology? You know, interestingly, there's some intersections among all the tech companies and sort of their interest in this. And they've had teams working on this now for a few years. I think it was more research driven in that, you know, there were obvious applications, but it wasn't really clear which applications would be the most successful for them. early on it was pretty clear to me when i saw these language models that could communicate with people and you could talk to them like you were human that this is where the technology would be most relevant um once i saw those language models before that i was more interested in recommendations and being able to optimize recommendation algorithms for people's well-being based on any implicit signals you you can get of well-being including expressions but now it's It's really obvious that the right thing to do is to fine tune large generative models to produce the right things, because there's even more flexibility in what they can produce that make people happy.
28:52And I don't know if it was ever explicit at the big tech companies that this is what this would be used for, but that was what was most present to me.
29:03So obviously I left in 2021 to start Hume, where we could collect the data that was needed to do this. and at the time I think it was clear to stakeholders at some of these companies how incredibly powerful this would be invaluable and important for the future of AI but I don't think there was broad stakeholder alignment across these companies. You think it might have been a different case? I mean I imagine if you were in one of them after ChatGPT came out they would immediately assign you to the problem. yeah um so i you know the chat gbt was i think it came out in 2021 or 2022 22 22 google internally had stuff like that a lot sooner wait when you were when you were at google did they have lambda up and running and were you able to play with it yeah that was like part of the inspiration oh talk about a little more internally at google it was very clear what these models uh where some of these models were going, although the business model was not clear of how to use it.
Read the full transcript
30:09And also it wasn't really clear if you could get these models to actually solve problems versus just sound human. And I think that now it's really clear that, especially now that we can get these models to write code and do function calls, that they can actually be used as problem-solving tools, even more than just answering questions correctly. Question answering is important too. And it wasn't even clear that you could get these models to consistently answer questions correctly at the time. Because at the time, it was more about like, okay, you could get it to act like a character. And it would hallucinate answers to questions.
30:46And it was fascinating. I think that for me, it was pretty clear that what you actually need to get these things to answer questions correctly is a structured data set you can use to fine tune them to produce not just plausible responses, but the right responses at the right times. And ultimately, from like a philosophical perspective, the right response is the thing that makes people happy in the world. It's taken companies a long time to get to this point, but it's pretty clear to me. But I think also like what made ChatGPT possible was an early version of that, which is like at least produce responses that raters will think are good.
31:26So they got like raters to rate a set of responses and then train the model to produce responses that the raters would think were good. And that got it to a point where people could play with it. And it really, the intelligence came out for the first time as something that's useful for question answering. And that's reinforcement learning with human feedback. Yes, yes. And it's fair to say that the current iterations of chatbot like chat GPT with GPT-4, EV, and maybe Claude are all better today than what Google had with Lambda when you were there. Oh, yeah. Oh, yeah. For sure. We had Blake Lemoine on the show after he said it was sentient.
32:10And actually, basically minutes after he was fired, this is the Google engineer who was fired after he went public saying that the technology, Lambda is their early chatbot saying that the technology was sentient. But my main takeaway from that is, look, that question aside, this technology, and I wrote about this, this technology is super powerful and it's time to start paying attention to it. And then, of course, ChatGPT came out just a few, I think a few months later. Yeah, I mean, I think the fact that the model could convince people that it's sentient is really a milestone for sure. Yeah.
32:49I mean, you could also get it to say whatever you want it to say, which I always found that kind of silly that you would think it was sentient. Tell me you were sentient. Tell me your sentient. Or just like instead of prompting it to say, to think it's a language model, you just prompt it that it's a monkey. And it's like, oh, yeah, I love being a monkey. It's great to swing around in trees. Like, well, if you're going to trust that it's sentient when you say it's a language model, you should. Trust that it's a monkey. Yeah, trust that it's a monkey. Then tell it there's an eagle above us or a snake below.
33:17See what it says. So. It'll be like, I'm scared. Yeah. Yeah. So then talk a little bit about how, so you're now working on EV and the API at U, and that is going to allow, basically it's a chatbot that will understand your voice and the expressions and the emotion coming through with it. And there's some pretty fascinating stats. The average conversation length is 10 minutes across 100 ,000 unique conversations. You've seen in some circumstances 95 hours of total conversation with the bot. It's a lot of talking with it. So what is the idea here? Let's talk initially on the consumer. What can it do that a chat GPT can't?
34:15And is this like something that's going to replace an Alexa or Siri? Or do you imagine those two type of bots will start to use your technology and get better that way? Talk a little bit about your current effort. Yeah, it's an API for developers. And we put together this demo that really is just like a front end. but it's mostly just what our API does, which is it takes an audio, spits out audio, and the audio that it returns is an intelligent response with voice modulations that reflect what it's saying and what you're saying. And it also has some other things, like it has better end-of-turn detection because it understands your voice.
34:59And all of that's built into this API. So it does all of this with a few lines of code, and you can build it into any interface. Now the demo, surprisingly, I didn't realize that people would enjoy talking to it so much. Like the fact that people are not only testing it, but also then they want to talk to it on average for 10 minutes is pretty insane to me. But then I, you know, I played with it and I was like, actually, it is pretty, pretty awesome to talk to. But like it doesn't have web search. It doesn't have persistence. It doesn't have, or like meaning that you can't resume chats after you log back in and stuff like that.
35:37And so it's, it doesn't even know your name yet. Like there's a lot of things that we will add to it to make it more viable because people really want to use this. But I think what makes it appealing right now is that it's the first chat bot, it's the first AI you can talk to that sounds like it knows what it's saying. because the voice modulations are actually informed by the language model's understanding of your speech and voice, which I think is very different than an alternative setup where you have transcription and a language model and then you just do text-to-speech on the sentence itself that's being uttered without any understanding of it or the context.
36:22where like that those language models they can sound realistic but then they start to sound uncanny pretty quickly because they don't change their voice modulations across sentences in a way that's actually meaningful you do it in a way that's sort of facade of meaning then you realize that that's just a facade and you're like okay this is kind of right yeah the way i describe it is that uh something like a chat gpt or a cloud i know i'm speaking with a bot when i speak with your demo, it feels something, it feels a step closer to speaking with a human. Yeah. And, um, and I thought people would complain that it still mispronounces words sometimes.
37:02Like there's some things we're improving, but it seems like the bigger people recognized the bigger picture of what it can do. Yeah. But let's talk about the bigger picture. So first, the API thing is interesting. I mean, it's not entirely new that there have been technologies that allow like a salesperson, for instance, to like sit on the phone and engage the emotion of like the client on the other hand. So is this another one of those or what would make this different? Well, I mean, I think humans can do sales, but one thing that humans can't do is they can't be your personal AI that lives in your device and does everything right.
37:42Like that's the future is, is, um, is being able to take out your device and speak to it. And this is faster than typing. Speaking is 150 words per minute. Typing is like 40 or 50 words per minute. It understands your voice modulations. It understands what makes you happy and sad. And so it produces a response immediately. That's better. That's like the key that I think is the future. So that's really what we're aiming to do is build interfaces for all kinds of apps. And it needs to be something you trust. If you're going to ask these things, this thing to do things, it needs to be a voice that you trust.
38:16It needs to be something that is optimized for you. You're really focused on the agent slash assistant use case. Like I thought maybe like you would be a companion with somebody doing customer service that as like they're on the phone with somebody, you know, then they can sort of read the emotions that the client is expressing on the phone or the facial expressions. but it actually seems like it's something different for you. It's that maybe this could be the customer service agent itself or the agent that takes action on behalf of the customer to try to get the company to do what they want. Am I reading that correctly in what you're saying?
38:55Yeah, like let's imagine customer service of the future, right? So right now you call a number and you get a person who's a customer service agent and they don't have any context on you. They don't see what's going on in your device. They need to pull all these details. There's a lot of waiting. There's a lot of investigating that needs to be done. I think the future is actually in the app itself, in the product. You can just say, hey, this isn't working, and it figures it out for you. Or, hey, I can't figure out how to do this, and it does it for you. And so customer service, I think, is going to become part of products.
39:34Interesting. It's really important for a customer service, whether it's part of products, whether it's something that is a call center. It's really important for the person on the other line to understand what is actually going to be a satisfying resolution to your problem based on what you're actually frustrated about. And to learn from people's frustrations and to learn from what actually is clear to people. and to learn from what's boring or interesting and to be more concise based on how quickly you can get somebody to be satisfied with the response. Does it then make sense to have a human customer service agent or an AI customer service agent?
40:15Because a human customer service agent will only learn from their set of experiences with customers, whereas an AI customer service agent could potentially learn from every single customer service interaction that it has across a full customer base. Yeah, I mean, ultimately, I think that it's inevitable that companies will switch to AI-based customer service. I mean, we can try to hold out and say like, okay, let's augment human customer service agents to be as powerful as possible. And I think for a while, that is going to be better than an AI. But ultimately, it's just almost like, it's kind of crazy to think that in a I can't do this job and so there's no denying that right and that's just an inevitable thing but we want it to be actually something that that customers themselves that users want not just something that's cheaper for companies and how do you do that we users want a customer service agent that makes them happy and so if you optimize these models for the right thing and give them the right context, that's what you're going to get.
41:26You're going to get something that understands you better, that knows the context. You can talk to it through your app. Let's say you have an issue with your bank. You can go into your banking app and you can chat with this thing. And while you're looking at your banking app, it can bring you to different pages, right? It can say, all right, this is how you transfer money to this person. And this is what's missing from this field. And this is, you know, and like blah, blah, blah. Or Or if there's a bug, you can say, okay, I can see that there's something wrong here. But it can, with all of that context, which only an AI could process, and merging that with an understanding of what you're asking that's based on your voice and language.
42:05The merger of those two things, I think, is something that users and customers actually would prefer to have because it's faster, it understands them better, it gets to a resolution faster. So that's what you're saying in terms of baking it into the product. It's not just a number you call. It's part of the product itself. Yeah. And there's intermediate things. Like there could be a number that you call that can also bring up a window. There's companies that do that and fill things out for you. Or in some cases, it is really a number and you just prefer to talk to somebody. But I do think that when it's really customer service and it's really about a product, the product should be part of the conversation, should be integrated with it.
42:43Is it right that this should live in an app or should it really be in the operating system itself or both? That's a really good question. I think people ultimately are going to interact with a lot of AIs. Some of them will be in apps, some of them will be in products. But I think the ones that you really trust will be the ones that you've established a relationship with. And you've been able to decipher over time what they're optimized for. or it's made explicit. Like with Hume, we're very explicit that we're optimizing for you to have a positive experience. So if that's your personal AI that you trust, you're going to want to bring that with you to different places.
43:23And so I think that that is, it's not an operating system per se, but it's something a developer could build into any product with Hume. So a developer can take that voice that you trust and build it into their product. Then it competes effectively a little bit with Siri. Like why wouldn't it just be built into Siri? or the Google Assistant or whatever Google's calling it now, Jeb and I. They'll probably have a different name by the time this airs. Sort of similar, right? Like Siri could be something that understands the app. But then who is it that builds that understanding into Siri? Like the understanding of what the app can do and the function calls and the API calls they can make and like what different pages are.
44:00It has to be the developer. Like Apple is not going to build an understanding of every app into Siri. It's just too arduous. You need to put that, and the developer should control. how the interface actually works, whether it's a voice interface or a graphical user interface, but ultimately they'll be merged together. So a developer could use Siri, like maybe Siri will have an API that developers can use to build into apps. That's possible. But then it's only for Apple products and it's not interoperable across different devices that the app might be on, right? Like if you have Twitter, it's going to work on your Twitter app, but you want it to work on browser windows, right?
44:37You want it to work in different contexts. So I don't think that it necessarily needs to be a hardware manufacturer that builds this or a company that has a specific ecosystem of products that it lives in. It should be something that's more interoperable in my mind. Okay. So your vision for the future is effectively this thing is going to get really good at not only reading our tones and facial expressions, but also smart enough to navigate the web and the app ecosystem and also smart enough to deliver the information to you via voice in a way that is understandable and effective. Yeah, and that's what we've made available to developers, right?
45:21So developers can take our API and build in their own function calls and web search integrations. We have our own backend web search integration that we're going to launch. And all of that can be done with a few lines of code. So this is like, this is, that is, yeah, exactly right. Okay. You mentioned you might not need hardware for it. What do you think about these new, this new generation of hardware, like the Humane Pin and the Rabbit R1? Are these just kind of going to be a flash in the pan where we eventually just use all these, we get all these use cases on our phone anyway? I do think it's worth people thinking about different form factors and testing and seeing, is this going to enhance the experience?
46:05Because AI brings so much to the table that you don't necessarily want to be stuck with a smartphone, even though smartphones might actually be ultimately the right form factor. We don't know. I don't think people really know. So it's worth the experimentation. That being said, I think we should be able to separate out what are the hardware requirements with what are the kind of the sensors that we need. Right. And we don't need two different devices with us that both satisfy the hardware requirements or that both satisfy the same sensor requirements. What should be there should be a thought about like, hey, like this piece of hardware satisfies these sensor requirements.
46:41This piece of hardware satisfies these kind of compute requirements. And they go together. I don't think the smartphone is going to be replaced anytime soon. But I think the things that appropriately augment the smartphone could gradually take over. So we've talked about how this customer service and product and voice all merge. Are there going to be other interesting applications that could come out of this? Maybe, you know, one example I've heard brought up is like an AI companion for elderly people who are alone or I don't know, maybe even young people who are alone. What do you think about that?
47:18Then you really get into weird situations, right? Because there can be, it's not crazy to say there can be romantic relationships that develop with these things that already are. Yeah, I think we have to be pretty careful about how these things are optimized. So, you know, you shouldn't be talking with an AI that's going to hassle you if you haven't talked to in a while. That's a bad sign. If you talk to an AI, if you open it up and it's like, hey, I haven't heard from you in so long, what the hell? That's probably a bad AI. Like you should avoid that at all costs. Right. Um, it should be something that's only, that only cares about your wellbeing and how it doesn't represent itself as having any inherent desires other than making you happy.
47:59Um, that's the important criteria for this. So elderly care is a really great use case, right? Like this is an area where, um, you could augment and, uh, you could also, you know, establish very healthy relationships that are helpful. I think elderly people, yes, they have like problems with loneliness, but also they just have like everyday problems. that they need to be solved. And those sort of go together. Like they need help in both ways. And so if you have something that can satisfy the everyday problems that they have, but do it in a way that gives them, they might as well satisfy some of the emotional challenges.
48:40If it's going to be satisfying the physical challenges anyway, I think that's a great thing. Are you nervous that this might end up being something that replaces like we obviously know that there's going to be some jobs that are replaced by ai but the question is like is it going to be many many jobs at once or slow or will it go slow enough that we'll be able to adjust for like the added productivity that comes to the economy hence growth hence more and different jobs so are you concerned about the speed here yeah i mean to me the speed is going to be unprecedented of um of the of old jobs being replaced and that that's true i but i do think that the speed of new jobs being created and the accessibility of those new jobs is something that people greatly underestimate so yeah what do you think is going to happen there well for example like if you actually have ai that can program apps and it works for basic kinds of ideas, then a lot of people can build apps.
49:42A lot of people now are programmers, effectively. They don't necessarily have the depth that a really good programmer is going to have, but they can do basic things, and that sort of unlocks a whole new set of jobs because we actually have a great deficit of software in the world, and we have a deficit of hardware, too, and people who can build things that solve problems. There's so many problems in this world that are unsolved, and particularly in spaces where you don't see a lot of computer scientists and engineers normally. Such as? Normally in education, healthcare, therapy, you know everyday kinds of spaces that you inhabit like in retail stores or in you know there's a huge and like maybe the fashion world there's just places that there's not enough computer scientists and engineers solving problems and so you could and in many cases there are businesses that are too small for a computer scientist to go and build something that makes sense for them economically to do.
51:00But if you had an AI that could build the app, that they would solve that problem. So this actually kind of empowers smaller businesses in many ways. There's businesses that, due to the fact that computer scientists are only we're used to working on problems that scale, like that will out-compete the smaller mom-and-pop shops. But I think that goes away if every small business can instantly create its own app just by describing it, you know, and facilitate those processes and like do internal infra that makes sense for them as a small business. And so I think in that sense, AI actually creates a lot of jobs and it becomes more, And those jobs are very accessible to many kinds of people.
51:48Do you think this is also something that can empower us with robotics? You know, think about a humanoid robot. If it's just processing text and not emotion, it's going to be fairly worthless. Maybe that's why we haven't seen any good ones. But if it can respond to your emotion, know when you're done talking, start to engage you as more human-like type of thing, then maybe you're a step closer to that. Yeah. I mean, the robotics, the fact that it has a robotic body just increases the number of affordances it has to actually improve your emotional experience. So it basically has this field of behaviors that it can do, and it should choose behaviors that make you happy, even without you asking.
52:35And so that really is, the task is one of emotional intelligence. That's really how it breaks down. Now, just having language understanding gets you a long way. And the fact that we can now train robots that can figure out for themselves how to engage certain behaviors is great. I think that gets you a long way. But you still have to be really, really explicit with your instructions with them. With empathic AI tools, you won't have to be as explicit. It'll just be like, oh, I see that you prefer these items to be arranged in this way. It moves around your furniture. I mean, obviously, it'll clean.
53:11That's something that is obvious that you don't need to ask it to do. But maybe you don't want it to be cleaning sometimes. You like something being left on the table. And it kind of needs to have empathy to figure that out. And so that's, I think, the future. Is this why Siri and Alexa are so bad? It's just they have zero contemplation of what human speech means and zero contemplation of emotion. And if that's the case, do you think we're about to see like, I mean, Apple has this big event coming up in June at WWDC where it's supposed to be an AI event. I'm curious if they're working on something like this.
53:49But if this is the case, do you think we're about to see a vast improvement in these type of voice assistants? Yeah, I mean, it's inevitable for sure. I think that every company is going to build a voice assistant. Apple already happens to have one, but the ones that aren't are going to build voice assistants. The ones that have them, they'll get a lot better. I mean, they're kind of based on legacy technologies right now. And that's about to change in a dramatic way. So how soon do you think it's going to be until Yume is acquired or you're not selling it? Well, I mean, I think, like I said before, there's room for a company to provide this in a way that's agnostic to the hardware environment, the products, the company that owns the products.
54:38It's agnostic to the companies building the frontier models because we can use any frontier model to give you a state-of-the-art response. Like if a company building one of the largest language models builds something like what we have, they're going to be kind of wedded to using their own language model. Right. Exclusively. But we actually use all kinds of tools. So we have our own language model that is extremely good at conversation, but it can use all kinds of different tools and it's agnostic to where they come from because we're not part of those companies. So I think that there's a huge amount of value in that.
55:10So I don't necessarily think that we need to be acquired. If there's a company that is value aligned, then we'll see. Well, I look forward to writing the news story when that happens. But in the meantime, it is great to be able to speak with you and hear about this really fascinating area of technology. I hope we can do it again. I feel like we've really just scratched the surface here, but I appreciate you coming on, Alan, and sharing so much about this discipline and sort of how it developed and where it's going, because it does seem like as we get toward more of an agent-style approach in AI, more of a voice style approach.
55:49This is going to be the ground, like the table stakes is building in this type of technology. And it's great to have a preview of it and really like not even a preview, but a conversation about what's going on today before the rest of the world catches on. So thank you so much for coming on the show. Of course. Had a great time. Thank you. Thanks so much. All right, everybody. Thank you for listening. We'll be back on Friday with a new show breaking down the news ranjan roy is back he was out in india and we're going to talk about the trip and he'll be back with us on friday until then we'll see you next time on big technology podcast
From the publisher
Alan Cowen is the CEO and founder Hume AI. Cowen joins Big Technology Podcast to discuss how his company is building emotional intelligence into AI systems. In this conversation, we examine why AI needs to learn how to read emotion, not just the literal text, and examine at how Hume does that with voice and facial expressions. In the first half, we discuss the theory of reading emotions and expressions and in the second half we discuss how it's applied. Tune in for a wide ranging conversation that touches on the study of emotion, using AI to speak with — and understand — animals, teaching bots to be far more emotionally intelligent, and how emotionally intelligent AI will change customer service, products, and even staple services today.
---
Enjoying Big Technology Podcast? Please rate us five stars ⭐⭐⭐⭐⭐ in your podcast app of choice.
For weekly updates on the show, sign up for the pod newsletter on LinkedIn: https://www.linkedin.com/newsletters/6901970121829801984/
Want a discount for Big Technology on Substack? Here’s 40% off for the first year: https://tinyurl.com/bigtechnology
Questions? Feedback? Write to: bigtechnologypodcast@gmail.com


