In short
Podcast Summary: This Week in Startups - Empathic AI and its Role in Understanding Human Emotions with Hume AI’s Alan Cowen | E1922
Episode Overview In this episode of *This Week in Startups*, host Jason Calacanis interviews Alan Cowen, CEO and Chief Scientist of Hume AI. They discuss the innovative technology Hume AI is developing to understand human emotions through AI. The conversation covers the Empathic Voice Interface (EVI), their Measurement API, and the implications of using AI in emotional understanding.
Episode Details
- Host: Jason Calacanis
- Guest: Alan Cowen, CEO and Chief Scientist at Hume AI
- Key Topics: Empathic AI, understanding emotions, emotional intelligence in technology, AI applications in customer service and therapy.
Key Highlights
Introduction to Hume AI (3:01)
- Mission: To optimize AI for human well-being by understanding emotions and their expressions.
- Focus: Moving beyond language alone, integrating vocal tone, facial expressions, and emotional states into AI responses.
Empathic Voice Interface (EVI) (6:08)
- Demo: Hume AI showcases EVI, demonstrating its capability to recognize and respond to human emotions in real-time.
- Functionality:
- Analyzes voice and facial expressions to discern emotions like sadness or frustration.
- Provides contextual responses based on emotional analysis.
Measurement API (16:20)
- Purpose: Allows developers to integrate emotional analysis into applications, enabling real-time emotion recognition and response.
- Capabilities: Offers a nuanced understanding of emotional states, producing metrics and feedback about user emotions.
Implications of Emotionally Intelligent AI (44:11)
- Positive Applications:
- Enhances customer service by guiding representatives on how to respond based on caller emotions.
- Potential use in healthcare for therapy support and tracking emotional well-being.
- Concerns:
- Risks of manipulation in marketing and politics if AI is used to sway opinions based on emotional cues.
- Ethical considerations in deploying AI technologies that influence behavior.
Key Concepts Discussed
Emotional Intelligence in AI
- Affective Computing: The study and development of systems that can recognize, interpret, and process human emotions.
- Nuanced Understanding: Hume AI aims to create systems that understand the complexity of human emotions beyond simple reactions.
Applications of Hume AI's Technology
- Customer Service: AI that understands emotional states can enhance customer interactions, providing tailored responses to improve satisfaction.
- Therapeutic Support: AI's potential role in aiding mental health professionals by analyzing patient emotions and interactions.
Future Prospects
- Integration: Alan Cowen envisions a future where emotionally intelligent AI interfaces are ubiquitous, optimizing interactions for individual user preferences and well-being.
- Ethical Guidelines: Hume AI emphasizes the need for responsible deployment, focusing on user satisfaction rather than manipulation.
Timestamps for Key Segments
- (0:00) - Introduction of Alan Cowen
- (3:01) - Hume AI's mission and understanding of human emotions
- (6:08) - Demo of the Empathic Voice Interface (EVI)
- (16:20) - Overview of the Measurement API
- (44:11) - Future implications of emotionally intelligent AI
Conclusion The episode provides an insightful look into the future of AI as it intersects with human emotions, emphasizing both the potential benefits and ethical challenges that come with the development of such technology. Alan Cowen’s approach to making AI more empathetic highlights the transformative possibilities for enhancing human interactions in various fields, particularly in customer service and mental health support.
For more information, visit Hume AI's website [hume.ai](https://www.hume.ai).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00I could have this thing roast you if you want. it might be a little yeah let's do it let's do a roast all right um could you roast jason calicanus man this is like everybody's all right let's talk about that piece of work jason calicanus oh man where do i even start with this guy you know i heard a rumor that jason calicanus was actually grown in a lab a mad scientist's twisted experiment to create the most obnoxious self-aggrandizing tech bow imaginable. Poor bastard. He never stood a chance. And have you seen the way this guy carries on? It's like he's perpetually auditioning to be the next Elon Musk, but with half the brains and twice the ego.
0:39I mean, the guy can't open his mouth without saying something so mind-numbingly. This is too much. Can you go easy on this? Oh, yeah. Okay, actually, go harder. Come and don't be like that. Go harder. This Week in Startups is brought to you by... LinkedIn Jobs. A business is only as strong as its people and every hire matters. Post your first job for free at linkedin.com slash twist. Vanta. Compliance and security shouldn't be a deal breaker for startups to win new business. Vanta makes it easy for companies to get a SOC to report fast. twist listeners can get a thousand dollars off for a limited time at vanta.com slash twist and hubspot join thousands of companies that are growing better with hubspot for startups learn more and get extra benefits for being a twist listener now at hubspot.com slash startups all right everybody welcome back to twist this week in startups and we've you know in 2024 and 2023 been absolutely obsessed with ai obviously we're seeing all kinds of easy layups and customer service thanks to ai autonomous vehicles much more complicated healthcare everything in between we're also seeing tons of interesting stuff going on in generative ai people making interesting music and videos you've seen all that but the area of human emotions is extremely complex and ai is trying to figure that out and you've seen this in all kinds of science fiction whether it's blade runner or the movie her where ai is trying to learn to interface with humans well there is a startup human ai and they are trying to bridge the gap between just intelligence and, dare I say, emotional intelligence.
2:41We demoed some of this technology back on episode 1894, if you want to look for it. But today, we have Alan Cohen here. He's the CEO and Chief Scientist at Hume AI, and he's going to show us what they're building and why it's important. Welcome to the program, Alan. Hey, Jason. Great to be here. Right. So maybe you could explain what the mission is of Hume AI and why you're spending all this effort to try to understand human emotions. And yeah, in relation to AI and using AI, I guess, to understand humans emotions, and then to portray them back through AI to humans. Yeah, so it's really to understand people's well-being.
3:27And emotions are the components of that. So when are you laughing? When are you sad? When are you in pain? When are you experiencing pleasure? And what we want to do is optimize for that. So our mission is to optimize AI for human well-being. Now, so much of what we express is in our voice, in our facial expression, and not in language. So that part of our expression was just ignored by AI for a long time. I mean, there is a field of affective computing. which I have a lot of experience in. I have over 40 papers in that area. But in terms of the generative models, they just were very far behind in understanding expressions.
4:02So what we've done at Hume is built models that understand expressions a lot better. And we've integrated those into large language models. So now these models understand beyond language, what's going on in the voice, what's going on in facial expression, and can learn from that. So they figure out what's making you frustrated, what's satisfying what's funny and they can actually adapt to that information and get better over time got it so what is the killer app of this genre or this vertical i have chat gpt i pay 20 bucks a month for it i frequently will use the interface where i talk to it but i don't believe that they are looking at my tone and i know i haven't turned my camera on i assume with your technology at some point, these AI assistants are going to have the camera on with permission, obviously, and hear my tone of voice and understand when I'm frustrated or explain to us what the killer app here is.
5:03Because, you know, we've heard of things like Gong and Rilla Voice, we have a small investment in. They study salespeople and how they behave on sales calls to coach them. That genre exists a bit. I'm curious what you perceive as the killer app here? So ChatGPT is text only. We think the future of AI is a voice app. The voice is four times faster than text. The problem is that when we're speaking, we expect the AI to understand not just what we're saying, but how we're saying it. So that's why this experience is clunky. It's just missing something. And so we've built that. We have a talking ChatGPT, basically.
5:40And we have an API that allows developers to build a voice interface into anything. So instead of using OpenAI, which would enable developers to build a chatbot interface into their product, they use us. And immediately they have a voice interface that does transcription, does language understanding, that speaks back. And what's different about it than anyone else is it understands what's in their voice besides language, and it uses that to form better words and better speech. Supposed to be a perfect time, I think, for a demo, so let's get to it. All right. You queued it up perfectly. sounds great I will share a few demos and what's the API called it's called EV empathic voice interface empathic voice interface okay like Lieutenant Deanna Troy on Star Trek Next Generation she was an empath yeah why don't I just demonstrate this this is like the demo that anyone can play with and then I'll show you the developer playground
6:41hey what's up Well, hello there. I am doing just fine. Thanks for asking. So what's got you feeling a bit down and amused at the same time? I am all ears if you want to chat about it. Well, my dog died this morning.
7:02Oh no, I am so sorry to hear about your dog. Losing a beloved pet is never easy. I know it must be really hard right now, but try to remember all the wonderful times you had - Aha! You really had me going there! Well played, you sneaky trickster. Aha! I should have known better than the fall for that one. I think I am still learning when it comes to your sense of humor, but I am relieved to hear your dog is doing just fine. Care to share what's really on your mind today? Well, can you tell what I'm expressing now? Oh, our demo's at capacity right now. Oh, there you go. So while you were doing that demo for people who were listening, and not watching.
7:44You can watch, of course, on YouTube, go to This Week in Startups on YouTube and you'll find it. It was showing each transcript in real time very, very quickly. And then it had the top three emotions and the percentage of those emotions. I think it was showing the top three every time. Is that correct? Yeah. So it shows more than just the top three. But actually, if you were to look at your raw data, you get back 48 different dimensions. So it's much more nuanced than what we're showing you there. Got it. And so in real time, you can see that you were sad when your dog died and etc. And then that person was showing sympathy for you.
8:21So all of that is being done through tone of voice, inflection, etc. Okay, let me cut to the chase right now because I know you're busy and everyone is hiring right now. And you know, it's a lot of competition for the best candidates, right? Every position counts, markets starting to come back, you need to get the perfect person. You want a bar raiser in your organization, somebody who will raise the bar for the entire team. And LinkedIn is giving you your first job posting for free to go find that bar raiser, linkedin.com slash twist. And if you want to build a great company, you're going to need a great team.
8:53It's as simple as that. LinkedIn jobs is here to make it quick. and easy to hire these elite team members. And I know it's crazy, right? LinkedIn has more than a billion users. We all watched this happen when it was tens of millions, then hundreds of millions, and now a billion people using the service. This means that you're going to get access to active and passive job seekers. Active job seekers, they're out there looking. Passive job seekers, they got a job, but it's not as good as the job you're offering them. So you want to get in front of both of those people. Maybe somebody got laid off, wasn't their fault, and they're an ideal candidate.
9:24Get that active job seeker. And LinkedIn also knows that small businesses are wearing so many hats right now and you might not have the time or resources to devote to hiring. So let LinkedIn make it automatic for you. Go post an open job role. You get that purple hiring ring on your profile. You start posting interesting content. You watch the qualified candidates. They just roll in. And guess what? First one's on us. Call to action. Very simple. LinkedIn.com slash TWIST. LinkedIn.com slash twist. That'll get you your first job posting for free on your boy j cal terms and conditions do apply what are the components in voice that you're studying is it the speed at which somebody speaks you know tone and and how did you train this thing on tone how does it know what sadness is versus you know melancholy versus quirky yeah we have all this data from millions of people around the world who are actually recording themselves while they're having interactions and also we're reacting to things and um and imitating things in some cases and so we use all that data to train our models and that means they're able to capture way more than just like tone rhythm like those are all basic things um but dimensions that you can't really describe in any other way except to say like this is kind of an angry dimension kind of has a growl to it kind of um tension in the voice where this is like an awe-inspired dimension or happy.
10:51And we get tons of different dimensions out of that. So every time we hear a word, we're getting more than 48 different dimensions of expression from that word. Our model is taking that in and our model is deciding how to respond. Our model is learning what these dimensions mean from tons and tons of data, people interacting. And it's saying, okay, this is something that means this person's frustrated, so I should apologize. It's something that means the person's confused or actually clarify and is figuring out what it should do to respond to somebody in different situations. How different is it per person?
11:24Like I might be a high energy guy from Brooklyn who's extroverted, who speaks a certain way and is, you know, you might be more introverted and soft spoken. So how does it know if J. Cal's like bombastic and joking, and you might be more thoughtful and introverted. Are our emotional emotions very similar or are they very disparate? I'm curious. So it has to learn that stuff. So we train in all these interactions, right? And so it's trying to figure out, the task is actually predicting the next expression. So it has to figure out like, is this next person going to laugh at what's said or like, are they going to be frustrated?
12:03And so it has to learn how you express your response to things in the course of doing that. and it's learning that in a generative way in a very ground up way. So by the time that we've trained this thing, it has to account for individual differences, for potentially cultural differences, for sentiments, and also just the average of all humans and like what humans respond to along with the distribution, if that makes sense. So like, what is it that humans find funny? What is it that humans find sad and all that? So like when I said my dog died, we can probably figure out this is a sad event. I'm going to be sympathetic, right?
12:36That's how it, and it figured out. How much of it is the words versus the tone of voice, or is it doing both of those things? It's doing both. Let me try to give this another shot so you can see that. Can you tell what I'm expressing right now? Whoa there. I can hear the frustration in your voice. It's an anger, determination, and distress. Like you're ready to tackle whatever's got you worked up. Can you tell what I'm expressing now? Hmm. I am picking up on some subtle shifts here. You sound a bit more relaxed now, though maybe still a tad bored or uneasy. But then I also hear a spark of amusement and even happiness.
13:18Like, you're pleasantly surprised by something. Am I on the right track there? I'm going to mute that bit. Got it. You're sounding a bit more at ease now with a hint of satisfaction. Anyway, you get the ideas. Yeah. So if that demo is designed to reflect back to you what emotions and things you're having, and then tweak it. So how long does it take for it to accurately understand a human? It's less than 500 milliseconds. As you can see, our API is experiencing some load right now. But generally speaking, we can get you back a response faster than any other API. and that's because we're able to detect when you're done speaking more accurately.
14:03So some of the other APIs, they have to do this dance between, can I jump in or is it going to interrupt the user? And so there's a little bit of a pause. But for us, because we understand the tone of voice, we can use that to figure out when the person's done speaking and then more accurately know when to step in. And so that enables us to respond a lot faster. So it doesn't need to talk to me and ask me 10 questions to understand my emotional state and how I might be uniquely different than another person. What about across cultures? Because do different cultures have, obviously we have different languages, but even putting aside language, does tone work across cultures?
14:39Do Koreans, Italians, and Americans all emote frustration the same way, anger the same way? Is it across cultures or does it require more subtlety? There's similarities and differences. We have a paper that just came out on this, but basically if you're speaking a different language, we need to train a new model for it. And it can be not a completely different model from scratch, but at least we need to fine tune on that language. That's what we find for most languages, especially for broadly different languages. Like all the Latin languages have similarities, but if you look across East Asian languages, things are pretty different.
15:18So, yeah. So suffice to say, yes, we do need to train things for each language. And this demo only works in English right now. So who's using the app? Like, let's take a look at the developer console. So you have that up there, the playground. Yeah. Who's using this now? And is it in production anywhere? And what are people using it for? Because there's plenty of models out there to give you answers and generate copy for you. I'm wondering if people are even up to this level of nuance in their products yet, or just trying to get correct answers, because accuracy seems to be a pretty paramount problem right now.
15:54Yeah, I mean, you might be interested in accuracy, but if you're using a voice interface, you need to get to the point fast, right? And so that's really what we're doing. And you can't, with these long, verbose responses from chatbots, first of all, those are very taxing on the brain to read, so that's not a good interface. but also you might have an accurate answer in there somewhere. It doesn't really matter if someone's not going to listen to a voice reading that out for three minutes, right? That's a good point. So we have a lot of developers lined up for this. We haven't released this API.
16:26By the time this comes out, we will have released it because we're releasing this on Wednesday. But so far, we have developers on this, which is our measurement API. Oh, wow. So you're on a webcam right now. I'm just going to describe it. and you're making funny faces right now uh you're surprised horror confusion sadness disappointed laughing and if i were to just say be completely calm and at ease your calmness just went up to 79 your concentration went up to 45 um and now if you um started thinking deeply about the meaning of the universe like why are we here like what is the purpose of life like why wake up and build this company every day says you're calm you're calm with existential wait is this a video or are you doing this right now alan i'm doing this right now you're doing it right now you're not following my instructions no give me your exit give me your most existential like i'm wondering about the meaning of life like why are we all here i want to see if it gets existential confusion there it is contemplation yeah contemplation well what's interesting about this is this would be great for coaching an actor because like happy's easy sad's easy if you go happy it's got joy amusement excitement great and if you were sad sad is disappointing and confusion maybe you're just not a good actor maybe you need to take acting lessons for these demos contemplation is a tough one I was trying to get you to have existential angst I was trying to pick something that's really hard to read well just think like should you even come to work is it all meaningless that's kind of depression it'd be sad yeah a little sad a little confusion boredom it's fascinating so this is just really getting your facial expression in real time.
18:31So if you were frustrated, the AI would know it and be like, huh, that wasn't the answer we were looking for. And so are people using this for therapy yet or like therapeutic coaching kind of things? Because that one seemed to be like, I got a lot of pitches for people who want to create AI therapists. And I'm like, hmm, that's a little dicey. I don't know if you should call it a therapist, but companion. are people using this for companionship? I do think that AI is going to be something that is your friend. And so it's not just like a niche application. I think generally speaking, we want an assistant that understands us.
19:10And there's tons of people working on that. I mean, there are people working on explicit therapy apps with Hume too. And actually, a lot of it's in training therapists. And it's getting them to... you know there's a delicate balance you don't really want to like comment too much on people's emotions but you want to ask the right questions and kind of get at it help them understand their emotions better and so there's a lot of that and there's also like therapist burnout doctor burnout there's a lot of health and wellness applications there's also tracking depression and stuff we work with clinical researchers who are running clinical trials and using Hume to track symptoms of depression and Parkinson's, just the symptoms.
19:51It's not like used for diagnosis because ultimately the doctor does that. But it's helping the doctor understand these things. So we have a lot of those applications. A lot of them, you know, those are interpersonal things. Like someone's talking to someone and we're already like the measurement APIs that we have are very good at extracting more data from that and helping people analyze it and helping people understand themselves and their patients, I guess. I mean, is it so if we have therapy on one side, you have the therapist who needs to present in a certain way to get people to open up. If you believe in that modality, if you believe in Western psychotherapy, there is something about pacing and aligning with the person matching their energy and getting them to open up so that they have some cathartic, you know, way of processing stuff.
20:44So people are using it to train therapists so that they don't have a goofy look on their face or they have the appropriate look that would elicit less suffering in their patients. Is that what I'm... Yeah, or like customer service reps, which is actually a kind of similar thing. Yeah, it's another form of therapy, actually. It essentially is, yeah. Yeah. But you know, that's requires somebody who's technical, who's maybe academic, maybe a researcher to take take these measures and make sense of them. Listen, a strong sales team can make all the difference for a B2B startup. But if you're going to hire sharks, you need to let them hunt and you can't slow them down with compliance hurdles like SOC 2.
21:26What is SOC 2? Well, any company that stores customer data in the cloud needs to be SOC 2 compliant. If you don't have your SOC 2 tight, your sales team can't close major deals. It's that simple. But thankfully, Vanta makes it really easy to get and renew your SOC 2 compliance. On average, Vanta customers are compliant in just two to four weeks. Without Vanta, it takes three to five months. Vanta can save you hundreds of hours of work and up to 85 % on compliance costs. And Vanta does more than just SOC 2. They also automate up to 90 % compliance for GDPR, HIPAA, and more. So here's your call to action.
21:58Stop slowing your sales team down and use Vanta. Get$1 ,000 off at vanta.com slash twist that's vanta.com slash twist for one thousand dollars off your sock too have you done this with poker players yet have you put poker players through this to see if they're lying or deceptive in a poker trade a lot of things and it does not it cannot tell poker players about thing you know at least professional poker players cannot tell i don't think the information's there i just don't think that with professional poker players that you can there's anything going on in their facial expression. What can you tell with people's facial expressions that we wouldn't know of?
22:35Some people have said you could tell a person's, if a person's sexuality, where a person's from, you could tell all kinds of interesting things that you wouldn't know. I think up to a level. Is that true or not? That's not really true. There's been a lot of pseudoscience in this area. Like most of the things that we can tell are things people want to communicate, which is good. We actually don't really care to impinge on people's things that they want to keep private. We're more interested in helping people communicate well and helping the AI understand what people want. And most of that's like they're overtly on the face.
23:14And for example, it even extends to things like, is the person done speaking? We're way better at understanding when they're done speaking because we can take into account facial expression versus just the language alone. And that's part of how our empathic voice interface is able to respond better. So you know when I say, this is the end of the sentence. Yes. Because of my facial expression, you get a quicker clue than audio only. Therefore, you can start speaking without interrupting me, which is what humans do with each other. Yeah. Like, imagine I'm speaking to you. And right now, it's clear to you.
23:52I just finished the sentence. It's clear to you I'm still speaking. And it's clear to you I'm going to say something again. But now I'm done. now i know i can speak right yeah which is what i do for a living on the podcasts is try to understand when people are done so that we can have the next person speak right like moderation is a is a difficult task um and customer support folks are using this already to understand how hot and bothered people are when they call the customer support line i assume to some extent yeah so kind of understanding is the customer having a good time bad time where are we kind of failing on customer service and which customer service reps are doing well or poorly and how do we train them to do better?
24:33How do we pull examples up of when they're not doing well so we can train them to do better? And, you know, there's a lot of AI going into customer support now. So some of our early design partners for this new API are people who want to take the automated customer support, make it a lot better, but still know when to include a human, escalated to a human exactly i mean that makes sense if the person's like this is incredibly frustrating you know and you start hearing the frustration go up and they're whatever united premier gold diamond status yeah you want to get them on the phone with somebody because it's you're you're starting to piss them off right yeah so understanding when that happens and how much of this is going to be used for security have you do you have any security applications coming because it's been well known like when you go to certain countries you know they'll ask you a couple of questions they try to read you do some human factoring and figure out if you're lying it's one of my favorite genres of television show is the people going through customs and they're trying to read if they're like sneaking into the country or sneaking things into the country are three-letter agencies using this technology yet to like analyze people as they come into buildings or we haven't been working with security yet not that we don't believe that that's a good application, but we're being a little bit more careful about how this is used and trying to make trying to make this as rigorous as possible, essentially.
25:58There's been a lot of providers of kind of like facial expression reading technology, who aren't very scientifically rigorous, make false promises, and then doesn't work. You're just like signaling out people for no reason, basically, which is not, you know, we want to we want to take, we want to be very careful to distance ourselves from those applications. I think there's a huge wealth of applications that are consumer focused, that are just helping people have a better time, a better experience with an application. I think that's a much bigger market. And then there's like the more niche things like security, lie detection, stuff like that, which...
26:34Yeah. Could you tell if somebody's acting suspicious based upon their facial expression or tone of voice, like they're knocking on your door, your ring doorbell answers and ring? is like, hey, this person is scoping the house out, you know, they could use your API to figure out the expression of people at the door. Yeah, I don't, I don't know. I feel like I've mixed feelings about that. I think in some cases, if you have enough data, you could start to do something in that direction. But if people are really good, if they're like, if people are criminals, they're pretty good at disguising their facial expressions, probably.
27:08Yes, they're sociopaths, they're criminals for a reason, right? So you would catch the neophyte criminals, the bad criminals, and the good criminals would just sell right by. You'd be like, yep, that person's 100 % trustworthy. And you'd let Hannibal Lecter into your house. Yeah, I can see how you need to be thoughtful about this. Yeah, we don't want to make promises that we can't keep, right? Well, I mean, yeah. So in that regard, let's go through the emotions that you feel you've nailed. And then tell us, what are the ones that are hard? I tried to do existential, and that was really hard, obviously.
27:40We're not there yet for existential crisis. yeah i mean um the way that you respond to expressions has more to do with almost subtle things that you can't name versus like overt emotions and it depends on the situation like if i'm using this like comedy thought um which i can prompt this is like a prompt that makes it a comedy thought basically um it'll respond to my emotions completely or my expressions completely differently than it than like a therapy use case so let's try it oh okay great whoa well well well look who decided to grace us with their presence tonight i don't worry i promise i'll keep it pg-13 for now you know i gotta say i am a little shocked that any of you fine folks decided to show up being seen on who in their right mind would willingly subject themselves to an evening of my particular brand of, shall we say, edgy humor.
28:36But hey, I am not one to judge. Can you make a joke about Alan Cowan, the CEO of Hume AI? Sure thing.
28:48Ooh, you want me to take a swing at the big cheese himself, eh? Well, twist my arm, why don't you? I'm going to say, I am not the biggest fan of that Alan Cowan character.
29:01talk about a guy who needs to have that massive ego of his deflated a bit am i right actually you know what i think i have the perfect way to do it um i say we round up a team of trained monkeys dress em up in business casual oh that's that's really offensive you've made me very upset whoa there partner it looks like i may have gone a little too far with that one huh uh sometimes i get a little carried away with the whole edgy comedian thing you know so like you can see in that case i don't want it to be that sympathetic like it's doing its thing and so you can determine what ai flavor you want to have interacting with folks you'd have one that's cheeky and playful you could have one that's super empathetic but in you know maybe not going over to patronizing but you could see this like if i'm calling the support line for united they might you know think i'm a new yorker who talks fast and i just want to get to the point or i could be from the south and into southern hospitality and it could take its time with me and ask me about the weather and how i'm doing and you know a little bit of chit chat some people like in the south i notice uh versus in new york where they're kind of get to the point.
Read the full transcript
30:19Let's move on. You can basically train your AI to have both modalities and dynamically switch between them. Exactly. So there's all this context, and then that kind of transforms the meanings of our expressions. So what an expression means and what to do with it really depends on all this other information that this model is taking into account. So it's not so much like detecting lies or detecting anxiety or detecting depression. It really depends on the context and we're able to integrate that into the model and then it's not just like these kind of canonical emotions like anger like there's a little bit of anger dimension in a joke you know anger and amusement and contempt maybe that makes it funny so it doesn't necessarily mean the person is expressing anger so to know what these expressions mean you really have to have the context you have to have the relationship that you're acting upon with your expressions and And that's what our AI does.
31:15So it's a little bit more nuanced than just detection. And these are all under what you studied, effective computing. Yeah, this is a specific school of computing that kind of bridges the psych department and the computer science and I think behavioral factors, industrial organizational psychology. Maybe you could give us a quick education on that. Yeah, so effective computing traditionally, it's the study of nonverbal expression, basically. So facial expression, the voice, body posture, and then, you know, most of the history of that is just labeling those things in a very predictive way. Now that we have generative models, we have large language models that can reason, it's really about reasoning about affect.
31:58And that's what we've introduced at Hume. So it's about understanding whether somebody is going to find something funny, whether somebody is going to find something confusing, and using expression along with language to come to those understandings. I would say historically that's not what affective computing has been, but now we've sort of pioneered this new form of affective computing that we're introducing to the world. Some of this was done. I know this was like a big thing that Minsky worked on at MIT, yeah? Did you go to MIT with him or did you? I went to Yale and then I went to UC Berkeley for my PhD.
32:34and I also worked at Google while I was at Berkeley and I helped start the affective computing team there. So I've been doing this for like 10 years. Minsky had, all the AI people had something to say about affect, right? But there really wasn't much that could be modeled at the time. Same with language, right? Things have come a long way. And I would say that there's affect in language and so the word affect has a little bit of misnomer. it's really computing with more than just language that we're doing we're computing with expression this is the way that expressions transform communication yeah because you you have a multimodal situation here you have the visual the facial expression you have audio and then you have the actual words right and so you're feeding all of those in at the same time to get the response and to understand the emotion yes and all this just contributes to accuracy.
33:31We can predict words better with expressions versus without. So like, if you look at the raw metrics that are used to train these large language models, we're doing better in terms of those raw metrics than models that just consider language alone. So this is like an intimate part of reasoning. And it's just part of human communication that we're now taking into account. It's not something that is niche. You know, I think people think about emotion and affect as these niche things that are important for therapy, important for comedians, important for like a few. But actually, this is something that's important for all conversation, important for any interaction with AI, just understanding a whole new modality of information that people use to converse with each other.
34:15Yeah, it's absolutely fascinating how quickly this has come together, because if we were sitting here two or three years ago, this just wouldn't be possible, would it? No. So, I mean, without large language models, without our measurement models, without the modifications that we've done to integrate those two things, like this was not possible at all. What has surprised you about what the AI understands and what your model understands and what has been either disappointing or challenging, you know, on this journey? So, yeah, that's interesting. I mean, linking together the language models and text-to-speech and transcription is something that other people are doing.
34:58But what we've sort of started to see emerge out of models that do all three that are linked together is that they have these emerging capabilities. And you start to see that in this interface where it's forming expressive speech. that's just like it feels different to me than if you just like link 11 labs and open ai and just have a talking chatbot like that just sounds it doesn't really sound like it's understanding you and this this is doing something a lot more nuanced do you understand what it's doing when you when it starts processing all this stuff and you feed it in do you actually know how it's coming to these conclusions or is it just sort of, you know, it's doing its best to figure it out and who knows?
35:47So yeah, we don't come in and tell it to respond to sadness with sympathy, but like it does, right? And it's sort of intuitive why that is. So I'm not going to say I don't understand that, but that's an emergent capability that we did not program in. And there's other things that it's doing that are more nuanced that we don't really have a handle on, except that we know what it's optimized for. Hey, everyone, you know, I'm obsessed with AI right now. And a fantastic report about how AI is going to change the game for startups has been released. It was published by our friends at HubSpot for startups.
36:18And it's great because they surveyed 1000 early stage founders to get you these insights. These are from the field. The report talks about AI tools and hacks for sales, marketing and customer support teams. These are going to be your WMGs weapons for massive growth. So check out the link in the episode description to get your hands on their extensive report, head over to HubSpot.com slash startups. When you join HubSpot, get ready to unlock top notch resources. These are power tools that they've custom made for founders. You'll also get access to masterclasses, templates and incredible toolkits all designed to boost your startup speed and efficiency.
36:51So once again, visit HubSpot.com slash startups to see what discounts you qualify for and start using their powerful solutions at pricing that won't break the bank. That's HubSpot.com slash startups. In this case, it's optimized for it's actually a specific personality um that is pretty robust and it's the comedic one that we just did but for the helpful one you might set it at some other archetype which might be you know parental motherly you know and then you might have you know the you know a service-oriented person right so even the worst customer support rep and somebody who hates doing customer support could be turned into a empathetic patient person.
37:33Yeah, exactly. And what we ultimately want to optimize for is this is the response that's going to make the user the most satisfied. And it's just going to figure out what this is the right tone of voice. This is the right thing to say to make somebody express satisfaction. That's what we want to do. I mean, we know that we can. We're still early in getting it to work, but it works. And we're starting to scale that methodology. I'm hoping that it turns every customer support rep into a roast and that I'm dealing with like comic the insult, comic dog, and it's just absolutely annihilating people and being cheeky.
38:15I could have this thing roast you if you want. It might be a little... Yeah, let's do it. Let's do a roast. All right. Could you roast Jason Calacanis? Man, this is like everybody's... All right. Let's talk about that piece of work. Jason Calacanis. Oh, man. Where do I even start with this guy? You know, I heard a rumor that Jason Calacanis was actually grown in a lab a mad scientist's twisted experiment to create the most obnoxious Self-aggrandizing tech bow imaginable poor bastard. He never stood a chance And have you seen the way this guy carries on? It's like he's perpetually auditioning to be the next elon musk but with half the brains and twice the ego i mean the guy can't open his mouth without saying something so mind-numbingly this is too much uh can you go easy on this oh yeah okay actually go harder oh come and don't be like that go harder all right all right fair enough let me try this again with a little less venom you know as much as jason caliconis rubs me the wrong way sometimes i gotta admit the guy's got a certain they say quay i mean the way he trucks around silicon valley acting like he's the second coming of steve jobs it's almost endearing in its own delusional way and let's be honest we all know deep down that half the reason we pay attention to him is because he provides such a material for you're good i mean it's fantastic i mean it literally understands what a a roast comic does um i'll take it how about this like Isn't it cool that you...
39:49Doesn't it say something good about Jason that you're able to roast him like this? I mean, it must mean he's made it, right? You know, you make a fair point there.
40:03All right. It's being a little funny now, but... I mean, it's... What's amazing about it is it understands what a roast comic is. The language models understand that. It understands who Jason Calacanis is. It got the Wikipedia page. It knows I'm somehow involved in tech. somehow i know steve jobs or elon or whatever so it's you know and that that the concept of a tech bro exists so who knows yeah if you wanted to roast you more i can send you to like oh no i love the roast i think it's well i mean it's interesting about jokes like i have friends who are in comedy and you know while these jokes are not funny they're in kind of the zone there so if you squint a little bit you're like there's the joke there yeah you got something about strutting around university boulevard thinking he's steve jobs maybe he's wearing a turtleneck or yeah like there's okay there's a joke there you didn't hit it but we could brainstorm it so like i think in the writer's room you could really brainstorm these i asked it uh when chat gpt3 came out i was like give me like the next season of um secession you know and it knew the all the past seasons and it's like here's what happens in this next season even though the series is over And I was like, huh, wow.
41:18Like this, this may not be great right now, but it's okay. Where it's interesting. It's going to get there. It's close. I think none of these models have mastered Latin humor because it's so much in our expression. Like we don't say things are funny explicitly because that would just make them not funny. So let me explain the joke to you. Exactly. That's when the joke didn't land. we have this new eval for humor and we're we're starting to push it basically we can optimize for laughter we can optimize for like what do people actually laugh at in millions of hours of conversation that's great and so that's how we're approaching these these kinds of problems so you could do a focus group where you had people watch curb your enthusiasm and you could say for a hundred people here's the funniest moments and for this demographic older people older men older women younger men teenagers gen x you could literally give you what jokes landed with each group yeah wow that's version one of this version two is like we have which we're doing now we have millions of hours of data and we analyze it just to see in general what's funny to people like across everything not in third group enthusiasm but like across everything every single thing in the world i can tell you like uh there's a great movie idiocracy and there's a amazing tv show have you seen idiocracy yeah great movie i mean it's so great but like everything's been reduced down to like its most basic thing like here's like a gel for you to eat like from a tube and the the hit show is ouch my balls which is just a compilation of somebody getting hit in the nuts over and over and over again just and it you know falls off of a roof lands on a fence falls off that gets hit by a crane with a big ball you know hitting him in his nuts ouch my i think it's called ouch nuts or something like that it's hilarious that's what it's been reduced down to somebody getting hit in the nuts we're hoping not to be too reductive but Yeah, maybe the AI will.
43:26I mean, you'll figure it out. You could literally crack humor. What language model did you build all this on? So we have our own language model and it calls other APIs. In this case, it's calling Claude. So Claude is providing the language response for some of the language responses, not all of them. We also have like a wrapper around Claude. It's not exactly a wrapper. It's our own language model that sort of integrates Claude into the speech to make it sound more conversational. And also like detects when you're done speaking and stuff. But we give Claude more data than just language. We give Claude like some of my tone of voice data, some additional data that we're getting through our APIs.
44:08So it's augmenting it as well. So eventually, what does the world look like? If you succeed with this and we're sitting here in five years and it's built into every iPhone and you figured out emotion perfectly, what do you think the world will look like? What are some highlights or, dare I say, dystopian, utopian sort of, what are the pros and cons of this technology going to be? So RM is utopian. We want to build a layer in between the application and these gigantic AI models that is decoding the user's intentions and preferences and relaying that information to the model. So that's what we have here, basically doing that with Cloud.
44:51and because we'll have facial expression and the voice, we're able to learn over time. We're able to build interfaces that understand you and what you want and are optimized for you. So suffice to say, basically it's going to be built into everything. It's going to be the universal interface that you use to interact with AI. That's the goal. And it's always going to be this AI that's optimized for your experience. So you can go to it and it knows basically what your preferences are, what makes you laugh, what makes you feel better, what you find to be a good explanation for things, your style of speech, how you write emails.
45:29It's going to know a lot of different things. Obviously, it's going to keep all this information very protected and private. Now, on the downside here, you could use this technology to say, I want to convince somebody subtly to vote you know this way politically i want to try to convince somebody that you know trump is amazing or biden's amazing or robert f kenny jr is the one so you could literally start creating robo calls or subtlety here using this emotion to try to sway people in politics um or towards ways of of being or thinking and we saw that happen with the youtube algorithm so how do you police how people use your system i saw you have ethical guidelines there and then obviously there's things that would be maybe r-rated or pg-13 and romance always comes up when people are doing it whether it's a blade runner or her so are people using this for romantic relationships and what's your take on allowing that and then also how do you think about influence big questions with TikTok today and your technology could really be used to influence people towards good and bad ends.
46:45Yeah, I think there's a pretty good way to operationalize the difference between when you're being manipulated by something that wants you to vote for a person or to buy something versus when you're dealing with an AI that's optimized for your own well-being. And that's what we try to do with our ethical guidelines. So we have this non-profit, but the human initiative that essentially tries to codify that principle and says, these are the ways that you can pursue these different applications so as to optimize people's well-being. It even has a bunch of ways that you can measure people's well-being that relies on a combination of what we're able to get through our API.
47:23So positive emotions, basically, over time. And also different kinds of self-report measures that we recommend gathering. So as long as the AI is optimized for your satisfaction, for your well-being, I think it's not manipulation. When you get an AI that's optimized for somebody else's objectives and using your emotions for that, then that can be manipulative, I think. And that goes for the romance case as well. If you're dealing with an AI girlfriend and it's ruining your life by forcing you... You're spending more time with it than you're spending with humans. and that's going to be a negative for you and it's going to show up in many ways as being negative for your well-being.
48:05That would show up in these measures. If it's optimized though for your well-being and you're having a good time with it and it's healthy and you're not spending more than X amount of time on it, maybe that's okay. What about trying to upsell me like, hey, you're in business class and would you like to be in first class? You're in economy. Would you like to go up to economy plus? And it uses sure technology to be really convincing about the value of that and upsell. How would you look at an upsell? That to me is... Ethical or not ethical? I think that's not ethical unless it's done in a very, very careful way.
48:44Basically, our guidelines don't allow that. But you could say... Guidelines don't allow an upsell? But humans do upsells all the time. Right. I think upsells are okay if the goal is to find the person who really will benefit from the upsell and only try to sell it to them. And you measure the effect of the upsell on people's well-being afterward. And you're like, okay, people actually benefit from this. I didn't sell this to somebody and then they regret it. So I think there's ways of doing it that are going to be fine. The problem is that if you just allow anything, if you allow people to optimize this for anything at all, then the potential for manipulation is pretty high.
49:25And I think this is true regardless of Hume. I think people are building these things that will be extremely persuasive. And Hume, ideally, will be providing the AI that responds and protects you. It's like, okay, I'm detecting... So you are very much in the camp of, hey, we have to be really thoughtful about how this technology is deployed. Yeah, but not in as much of a paternalistic way. Like, I think that technology can have a sense of humor, and it's okay if it offends some people, and it doesn't need to be politically correct all the time. But what I care about is, like, is this good for people?
49:57That's, like, at the end of the day. And so we have our ways of measuring well-being in order to optimize for that objective and not be paternalistic, basically. Yeah. But, I mean, at the end of the day, this is so powerful. It will be more powerful than just watching videos on YouTube because it's customized to an individual. so that ben shapiro or rachel maddow and pick whichever side of the political spectrum you're on you know those people are trying to convince you of their position and their interpretation of the world this to me is even more bespoke and customized to individuals so if you showed even a little propensity towards some of their viewpoints it could really whether it's the language model or the emotion but the combination of them you know the same way people were complaining like oh people go into the intellectual dark web on youtube i don't know if you heard about that like you you see a joe rogan then you get a jordan peterson you wind up on an alex jones and the next thing you know you're like some white supremacist or something is the claim um but yeah media does influence people and it is a stepping stone from one to the next to the next You may start out with somebody like Sam Harris, just intellectually rigorous, etc.
51:12And then all of a sudden, you wind up at Alex Jones as the complaint for many parents. But this would facilitate that, wouldn't it? Massively? I think you'd get that when you optimize for engagement. And so, to some extent, TikTok doesn't have this data, but it's still incredibly good at doing that. I think where this data helps you the most is in taking into account people's user satisfaction, their well-being, their mental health, all those things. So TikTok took this stuff into account and it was doing it in a way that was consistent with our guidelines, let's say. Then it would be using that data to optimize for people's well-being over time instead of engagement.
51:50And so, you know, you'd realize that if you throw people down the slippery slope of getting more and more extreme viewpoints, which is what happens today, because they're engaging and they're offensive at first and you want to argue. Like, if people who go down the slippery slope end up kind of isolated and it affects their social relationships, it affects, they start to get angry. This is not good for people's well-being. So the technology can look at that at the individual level, at the societal level, it can look at the health of all the people using a technology and say, hey, there's a collective impact of this.
52:29So that's the road we want to go down is being able to measure long term, is this affecting people positively? And you really need expressive behavior to look at that data. Like there's no other, like language alone is not going to get you there, basically. Yeah, I mean, you're going to be facing a real uphill battle because the marketers want this software. your top customers i predict will be marketers who want me to try zin or whatever those pouches are that people are putting there and you know i'm in texas right now and like everybody's putting these zin pouches in or whatever and i'm just like that can't be good for you and they're like want to try it like marketers love this kind of stuff like maybe the pitch to me is like be where you know hey it's performance and you know it's just nicotine it's like caffeine you drink caffeine you should try this and but for other people it might be you know hey you're the cool kid so it's uh you're going to be in a really interesting position as a provider of an api that i think a lot of the marketers are going to want to use this to try to convince people to do things that maybe it's unclear if it's actually good for them like hey you should you should gamble on sports right like there's marketing going on like crazy and if i'm a marketer, man, this is for me, the holy grail.
53:44Yeah, I think that we want to connect more with the end user and show the end user that we're optimizing for their interests and have that be the selling point rather than connect with the people selling to them. But I know I hear you. I think that's a real concern. But on the flip side, if we just optimized for people to buy things, or let's say we just optimized for engagement, you reach a certain point where it becomes so negative for people that regulators have to step in. And you kind of start to see it with TikTok, for example. Kids are spending six hours a day on TikTok. And if they made it any more addictive than it already is, parents would step in, regulators would step in, they'd be like, this is actually bad for our whole society.
54:27So at the end of the day, it's not necessarily good for our business. Yeah, which is what's happening with TikTok right now. As we speak, I think parents are getting the message like, this is too addicting for adults and kids. and the idea that like media is not influential is so naive like when people are like yeah you know media doesn't have an impact it's like are you sure like all studies show that media is one of the and video specifically is one of the most convincing mediums of all time and all human existence if you want to manipulate somebody video is the way to go and then customize video with, you know, that is matched to it is like 10x that.
55:04So you have like something here that I think is incredibly powerful. And the fact that you're being thoughtful about it makes me feel great. I think it's awesome that you're taking a measured approach to this. I wish you great success with it. If you want to learn more or try it, how do they get into the developer sandbox and play with this? And who are you looking to work with? Yeah, go to hume.ai. You can sign up. We'll be releasing access to our voice API hopefully before this episode comes out. And we have some closer design partners who we're working with to improve things as well. So please feel free to sign up and you can start using our API today, like our existing measurement API.
55:45I think it's absolutely fantastic what you're building. And I like the fact that you're super thoughtful about it, Alan, and I wish you great success with it. And we'll see you all next time on This Week in Startups. Bye-bye.
From the publisher
This Week in Startups is brought to you by…
LinkedIn Jobs. A business is only as strong as its people, and every hire matters. Go to LinkedIn.com/TWIST to post your first job for free. Terms and conditions apply.
Vanta. Compliance and security shouldn't be a deal-breaker for startups to win new business. Vanta makes it easy for companies to get a SOC 2 report fast. TWiST listeners can get $1,000 off for a limited time at http://www.vanta.com/twist
Hubspot for Startups. Join thousands of companies that are growing better with HubSpot for Startups. Learn more and get extra benefits for being a TWiST listener now at https://www.hubspot.com/startups
*
Todays show:
Hume AI’s Alan Cowen joins Jason to demo Hume AI’s Empathic Voice Interface (6:08), Measurement API (16:20), and discuss the future implications of this tech, both positive and negative (44:14).
*
Timestamps:
(0:00) Hume AI’s Alan Cowen joins Jason
(3:01) Hume AI and the role of AI in understanding human emotions
(6:08) Hume AI’s Empathic Voice Interface (EVI) and its responsiveness to human emotions
(8:27) LinkedIn Jobs - Post your first job for free at https://linkedin.com/twist
(9:55) The components in speech that Hume AI studies and its application across different cultures
(16:20) Hume AI’s Measurement API and its design for real-time emotion and expression analysis
(21:16) Vanta - Get $1000 off your SOC 2 at http://www.vanta.com/twist
(22:07) What AI can reveal about a person based on their expressions
(24:12) The impact on customer service and security sectors
(27:24) Hume AI’s comedy bot and emotional detection capabilities
(36:08) Hubspot for Startups - Learn more and get extra benefits for being a TWiST listener now at https://www.hubspot.com/startups. Also, be sure to visit https://bit.ly/hubspot-ai-report
(37:01) Hume AI’s comedy bot / roast functionality
(44:11) The future implications, both positive and negative, of emotionally intelligent AI on society
.*
Check out Hume AI: https://www.hume.ai
*
Follow Alan:
X: https://twitter.com/alancowen
LinkedIn: https://www.linkedin.com/in/alan-cowen
*
Subscribe to This Week in Startups on Apple: https://rb.gy/v19fcp
*
Follow Jason:
LinkedIn: https://www.linkedin.com/in/jasoncalacanis
*
Thank you to our partners:
(8:27) LinkedIn Jobs - Go to https://linkedIn.com/angel and post your first job for free.
(21:16) Vanta - Get $1000 off your SOC 2 at http://www.vanta.com/twist
(36:08) Hubspot for Startups - ****Learn more and get extra benefits for being a TWiST listener now at https://www.hubspot.com/startups. Also, be sure to visit https://bit.ly/hubspot-ai-report
*
Great 2023 interviews: Steve Huffman, Brian Chesky, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarland
*
Check out Jason’s suite of newsletters: https://substack.com/@calacanis
*
Follow TWiST:
Substack: https://twistartups.substack.com
Twitter: https://twitter.com/TWiStartups
YouTube: https://www.youtube.com/thisweekin
Instagram: https://www.instagram.com/thisweekinstartups
TikTok: https://www.tiktok.com/@thisweekinstartups
*
Subscribe to the Founder University Podcast: https://www.founder.university/podcast




