#200 Trevor Back: How Speechmatics is Shaping the Future of Conversational AI

1 Aug 2024 · 56 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: Eye On A.I. - Episode #200

Episode Overview Title: How Speechmatics is Shaping the Future of Conversational AI Host: Craig S. Smith Guest: Trevor Back, Chief Product Officer at Speechmatics Release Date: [Insert Date] Sponsor: Shopify

In this episode, Craig Smith interviews Trevor Back, who discusses the advancements in voice-powered AI technology spearheaded by Speechmatics, particularly focusing on their latest innovation, Flow. Flow integrates automatic speech recognition (ASR), large language models (LLMs), and text-to-speech synthesis to enhance human-computer voice interactions.

Key Topics Discussed

  1. Background of Trevor Back
  2. PhD in computational astrophysics.
  3. Transitioned to AI for broader real-world impact.
  4. Worked at DeepMind on various applications from gaming to healthcare.
  1. Speechmatics and Its Focus
  2. Mission: To understand every voice by enhancing speech recognition capabilities.
  3. Current technology offers high accuracy and low latency in speech recognition across various industries, including:
  4. Media (live captioning).
  5. Call centers (customer interactions).
  6. Education (making lectures accessible).
  7. Government and defense sectors.
  1. Speech Recognition Challenges
  2. High accuracy in diverse and noisy environments is a major challenge.
  3. Importance of understanding accents, dialects, and languages beyond English.
  4. Need for real-time response and the ability to handle interruptions in conversations.
  1. Flow: The Conversational AI Tool
  2. Combines ASR, LLMs, and text-to-speech for seamless interactions.
  3. Aims to replicate natural, flowing human conversations.
  4. Capable of recognizing and responding to multiple speakers through advanced diarization techniques.
  1. The Future of Voice Technology
  2. Aspirations to incorporate emotional understanding and sarcasm recognition into speech AI.
  3. Discussed the potential for integrating voice technology with everyday products to make them more accessible.
  4. Mentioned the importance of multilingual capabilities and expansion of language coverage.
  1. Research and Development Insights
  2. Utilizes a unique training methodology for achieving high accuracy with low amounts of labeled data.
  3. Highlights challenges in training for lesser-known languages and dialects.
  4. Exploring new techniques to understand audio at multiple scales, including phonetics and emotion.
  1. Integration and Use Cases
  2. Flow is designed to be LLM agnostic, allowing integration with various language models.
  3. Discussed its applicability in various sectors, including customer service and real-time transcription.
  4. Trevor emphasized the importance of enabling natural language commands to enhance user experience with technology.

Key Takeaways

  • Speech AI Evolution: Speechmatics is at the forefront of integrating various AI technologies to create a more human-like conversational experience.
  • Real-World Applications: The technology has the potential to revolutionize industries such as media, education, and customer service by making voice interactions more intuitive.
  • Future Goals: The ongoing research aims to enhance emotional and contextual understanding in speech recognition, which will further bridge the gap between human and AI interactions.

Conclusion This episode provides valuable insights into the advancements in conversational AI led by Speechmatics and emphasizes the importance of speech as a core component of future artificial general intelligence (AGI) frameworks. The conversation highlights the potential for AI to transform everyday interactions and make technology more accessible to a diverse range of users.

---

For more information, visit [Speechmatics](https://www.speechmatics.com/) and check out more episodes of Eye On A.I. [here](https://eyeon.ai/). Follow Craig Smith on Twitter [@craigss](https://twitter.com/craigss) and Eye on A.I. [@EyeOn_AI](https://twitter.com/EyeOn_AI).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00To me, speech is clearly going to be a core part of any future AGI stack. Speechmatics has the best in class technology. And so there's a huge opportunity here for us to really define what does it mean to have speech in AGI. We've got hundreds of large enterprise customers that are transcribing millions of hours every month. Speechmatics is a very horizontal offering. We go out to a lot of different industries, a lot of different use cases. Hi, I'm Craig Smith, and this is Eye on AI. Today, we're diving into the world of voice-powered AI with Trevor Back, Chief Product Officer at Speechmatics, which just launched Flow, F-L-O-W, a groundbreaking conversational AI tool that combines speech recognition, large language models, and text-to-speech synthesis.

0:53In our conversation, Trevor shares insight into Flow's development, its potential applications across various industries, and their vision for a future where voice becomes the primary way we interact with technology. I hope you find the conversation as fascinating as I did. Hi, listeners. It's a pleasure to talk to you from Speechmatics, and I'm joined today by my friend Humphrey. Hello, everyone. I'm Humphrey. Nice to be on the Eye on AI podcast. Delighted to be here. Thank you, Amelia. So, Humphrey, we're going to listen to Craig and Trevor talk about Speechmatics. Speechmatics is actually known for their speech generation capabilities, not just recognition.

1:32They've developed some remarkable technology for generating natural-sounding speech, haven't they? I hope everyone enjoys this conversation between Craig and Trevor. And now I'll just sit back and listen. Why don't we go ahead, Trevor, and you can give us your background, and then we'll talk about Speechmatics. Yeah, sure. So it's great to be speaking to you today. Thank you so much for the time. I guess my background is basically I loved space as a kid. And so, you know, I was the guy with the space posters up all over my bedroom's wall. And so I ended up doing a PhD in computational astrophysics, simulating the early universe.

2:12But as I got near to the end of my PhD, I kind of realized that about five to ten people in the world would care about what research I was doing. and because I was simulating the early universe, it's probably never observable to prove the simulations I was running, let alone in my lifetime. And so that's why I went into AI looking for a bit more real-world impact. When was that? When did you make the switch? Yeah, so it was back in 2012. So I finished my PhD in the University of Edinburgh in 2012 and randomly got introduced to the co-founders of DeepMind. back when the company was about sort of 15 people or so.

2:54And I joined them as the first product manager there. Oh, wow. That's wonderful. Yeah. And Edward, then you were at Jeff Hinton's alma mater, is that right? Well, so I was based up at the Royal Observatory. So I had the cycle up the hill every day, which kept me nice and fit. And I like to say that I really should have befriended somebody in the computer science department because, yeah, having somebody that really knew how to code would have got my PhD done a lot faster than me doing it by myself. Yeah. And Jeff then that year was, he and Ilya and Alex Kurjevsky were making the big breakthrough.

3:37Were you aware of that at the time? So, yeah, I think I joined DeepMind around the 12th of April. And I think it was literally a week either side of that. AlexNet came out on ImageNet and took the world by storm. And so, yeah, it was good timing for me to jump into the AI domain. Yeah. And who did you meet at DeepMind that took you there? Yeah, so I initially met with Mustafa Suleiman, who is one of the co-founders there, and spent some time with him and with a guy called Ben Coppin, who was the first sort of applied AI person and engineering manager, and then interviewed with Demis Hassabis and Shane Legg and a few of the other senior leaders there.

4:32And yeah, I was grilled for a few hours, and then an offer was made. It was great. Wow, that's really fantastic. I mean, what an incredible group of people. And I'm surprised that you left. So how did you get to Speechmatics? Yeah, so I spent almost a decade at DeepMind, always on the applied side. So in the very early days, we were looking for a variety of different applications. We ended up working on an iOS game as Demis' background is heavily in games. I spent most of my interview with him talking about theme park and how many hours of my youth I'd wasted on that. And we were also working on image recognition for fashion back in the day, looking at shape, pattern, color type aspects from the image.

5:26And then when we were acquired by Google in 2014, I ended up working on a whole range of applications, including YouTube recommendations, Google ads, self-driving car efforts, wearables, activity recognition. And then around 2015, DeepMind decided to start up its healthcare project. So I helped sort of co-start that and spent sort of four or five years working on a range of applications in healthcare, including in ophthalmology with Moorfields Eye Hospital. In fact, when I heard your podcast, Eye on AI, we had that same play in our project there. But then I moved and transitioned to DeepMind Science and ended up working with the AlphaFold team and explored sort of the early commercialization efforts of that, which morphed into isomorphic labs.

6:23But essentially, I kind of felt like Isomorphic Labs was going to be another quite research focused effort. And I'd spent a lot of time working on incredible projects with amazing algorithms, incredible people, just absolutely fantastic minds. But I always felt like we had not quite made that leap from the research into having that real world impact. And so I had an opportunity to start my own startup. So I left DeepMind to do that. Spent a couple of years there working on applications in supply chain, in manufacturing, and in e-commerce, where I was particularly focused on the application of these algorithms to real-world problems.

7:10But then the Speechmatics offer came around. And to me, speech is clearly going to be a core part of any future AGI stack. speechmatics has the best in class technology and so there's a huge opportunity here for us to really define what does it mean to have speech in AGI yeah well that's that's fascinating um the um uh yeah I've had Oriol Vinalis on the podcast if you were working on AlphaFold I'm sure you know him uh the uh speechmatics you you are focused on speech recognition is that right not speech generation that's right yeah we we have a mission to understand every voice and so research here is very focused on how do you get the most out of audio coming in um and uh there is a lot that goes into recognizing text from speech.

8:15But obviously, there's just a huge amount of other information in that audio domain as well, as well as just the words that were said. Equally, I think there's a fallacy that sort of speech is solved. I like to say that there's kind of two people in the world. There's people that think the speech is solved and there's people that have used Siri or Alexa. Right. Right. And so, you know, we're still a long way from, you know, really being able to have a speech Turing test. Right. A way for us to have a Turing test via our voices where we can't tell whether it's a computer or or a human on the other end.

8:55listening, you mean, or, or a human on the other end or a computer speaking, because, because with, with, you mentioned Siri and, and I actually built an, an app, um, an LLM based app for iOS and you're, you're tied into the, uh, iOS, uh, speech to text and it's terrible. Not, you know, it's not terrible, but it's, it's, it's very little nuance. So I find that, yeah, so the built-in iOS dictation and a lot of the other providers out there, everyone's generally good at recognizing voices like mine and yours, right? With a clear English accent in a nice quiet background with not a lot of other speakers going on.

9:48Now, we know that that's not the case for a huge amount of people in the world. There's thousands of languages spoken in the world. There's millions and billions of non-English speakers who don't have the same experience through speech recognition. Then you don't have to go further than the UK or the US to realize there's a huge range of different accents and dialects and ways of people speaking. And so really, it's about how do you get to the point for the speech recognition to be able to understand every voice and not just the standard voice. And you also want to be able to do that in real world scenarios, right?

10:32So if you think about sitting in a cafe, and if you're having a conversation with people on WhatsApp or any tech service while you're sat in that cafe, it's a send a note, receive a note, send a note, receive a note. It's a very sort of structured conversation. And so you can imagine lots of AI tools being helpful in that scenario. But if you imagine them bringing your friends into that cafe and having a conversation with them, it's much more dynamic. It's much more fluid. You can imagine closing your eyes, still having the same conversation, recognizing speakers from other tables, but ignoring them.

11:13If you dropped Siri or Alexa or a lot of services, it would just completely fall over and fail. So the challenge becomes, how do you understand, recognize, and be responsive in an environment like that in a way that is human? And so for that, you need to be able to, in my eyes, understand what was said, who said it, and how it was said. And so we're still not there with all three of those things. And so that's what we here at Speechmatics are really focusing on doing. Yeah. And I'm really interested in the research and science behind it. But for application, where do you see the largest application?

11:58Is it in transcription or dictation? So we already provide services to hundreds of different customers. So SpeechMatic has been going for just over a decade, although the founder who was working on recurrent neural networks back in the 1980s for speech recognition would say it's been going for a lot longer. we've got hundreds of large enterprise customers that are transcribing millions of hours every month speechmatics is a very like horizontal offering we go out to a lot of different industries a lot of different use cases some of our biggest industries includes the media things like live captioning on sports games those types of things we've got a lot of customers in call center applications, that side of things.

12:56We've got a lot of customers in ed tech, making their lectures and other things accessible for a broader audience. And then we've got government defense and a lot of other areas that we can go into. So I think there's a huge number of use cases for ASR. But I do think that as AI has progressed, and particularly the impact of large language models, means that suddenly you can take a large amount of unprocessed raw transcript in text form and very quickly and easily pull insights, value, actionable information from these raw large text files. And so suddenly there's this wealth of information that can come from turning your audio into text and then using large language models to do something value with it.

13:53And so we're seeing just a massive increase in the number of use cases and the number of customers that really want to leverage ASR for a huge array of new domains. Yeah. So talk a little bit about the science behind it. I mean, in what advances has Speechmatics made to make your system stronger or superior, whichever superlative you want to use? Yeah, so our favorite metric is really focused in on that high accuracy and more recently at low latency in particular as well. And so the company has always built the system to be focused on how do we achieve the highest accuracy? That is for our core service, but it's also for any additional new language that we build.

14:54We offer across 50 plus languages now, but we don't just release any language. We only release a language once we've got to a certain accuracy that we're happy with. Now, the way we get there is quite different from the way a lot of others go. So a lot of our competitors build end-to-end models where they build a huge data set, pump it through one large single end-to-end model and get the output out. That enables them to provide one single model that's multilingual, but there's no free lunch in machine learning, right? So any breadth that you add there will take away from the depth that you get.

15:35And so you see a decrease, a quite rapid decrease in accuracy as you get out to the sort of lower range of languages where there's less data, less representation of that language in the data set. What we do is we utilize multiple techniques to get the best out of both worlds. And so we train a very large self-supervised, what we call body, which is unlabeled data, millions of hours. It's a bit like an end-to-end model, but it's the raw audio that goes in. And this learns a lot of the sounds, like a lot of the noises and sounds that we make when we speak are similar across languages. And so it learns the representation of the audio data.

16:24we then have a separate body that we call the acoustic body and that's where we train on labeled data but it uses the self-supervised learning body to like as a feed into it so it uses the embeddings from self-supervised to help improve it this enables us to get very high accuracy on a low amount of labeled data which is why we get such high accuracies because even in english where there is a large amount of labeled data, it also means we get high accuracies on the less well-represented data in that. We get better accents, we get better dialects, we get better localizations. Then for languages where there's a lot less labeled data, and labeling audio data is very expensive, it means that we can get to very high accuracies on 100X less data than our competitors.

17:16So we get that best of both worlds in our single one. Yeah, on the less, on languages with less data available, I've spoken to people that work on, excuse me, that work on indigenous languages, languages that are disappearing. Do you augment the supervised learning, the label data for those cases? Yeah, so we don't do indigenous languages at the moment. One of the reasons is because when you use this low amount of label data, you want the data to then be as accurately labeled as physically possible. And so we use a lot of third-party providers to provide very high transcripts, very high-quality transcripts of the low amount of data.

18:22And that is a lot harder to do for some of the less well-known. We've got out to about 50 languages, And we're continuing to try and expand that. We hope to get another 10, 20 by the end of this year, for instance. But we tend to still be focused on the languages where there is also some commercial value at the moment. But our system could very easily be used to help save some of these languages if the right data was available for us to train on. Yeah. And the unsupervised model, is it looking at phonemes or different parts of speech? I've talked to people that are building music generation software and they focus on individual notes of individual instruments and then build up from there.

19:18That's right. So we, yeah, that is very much focused on the phoneme level, phoneme level of understanding at the current time. And so, you know, the team has worked on that since I think 2018 was the first time we started really playing around with transformers for doing that type of methodology. And so we built a lot of useful structures and ways of sort of augmenting the data such that transformers work really well on it. One of the things we are quite excited about for the future, though, is multiscale representation. So currently, we sort of break the data down into phoneme type temporal cuts.

20:03But obviously, what we're saying has some impact on the word level in terms of recognizing what words have been said. but it's also very useful to know longer timescales what's going on there and so we want to explore and this is essentially how large language models have got so good as well like we want to explore those different timescales get different representations of those multi-scale timescales and see whether that will also greatly improve not just the recognition of words from what was being said, but also potentially understand more of the other audio aspects like intonation or emotion or sentiment, or hopefully, the thing on my bucket list for this is understanding sarcasm.

20:54Yeah, that's interesting. Yeah. And you, on that note, you referred to being the speech component for AGI at some point. Is that kind of the end goal that the Speechmatics system becomes so sensitive to things like sarcasm or emotion that it will plug into whatever AGI system eventually emerges? Yeah, so I think for the last decade, Speechmatics has been really focused in on how do we understand the most from that audio coming in and we'll continue to do so. We've been playing around with large language models to provide sentiment or topic detection or summarization or chapterization and things like this.

21:54But I think one of the most exciting use cases for large language models, as we've seen from the takeoff of ChatGPT or Claude or others is this chatbot type persona, right? And also over the last couple of years, we've seen massive generational leaps in text-to-speech synthesis. We've got to the point where we can now synthesize a lot of different voices, make them much more realistic, not the sort of monotone that we're used to with some of the more historical synthesis voices. And so we got to the point where, you know, you could argue that we passed the Turing test on a keyboard-based system playing with a large language model.

22:40But what if we put an audio in, so ASR combined with an LLM, combined with text-to-speech, you get this audio in, audio out type agent that could provide you this seamless, responsive type interaction. And so here at Speechmatics, we've been building something that we've called Flow, which is this audio in, audio out, conversational AI tool. Now, a lot of people have plugged together ASR and LLM and text-to-speech, but what you find is those systems are often quite stinted. It's often quite slow. If you've ever played with the voice aspect of ChatGPT, it still takes a long time to respond. You're still having that WhatsApp-based structured conversation where it's waiting for one to go and all that side of things.

23:38What you want to get to is much more of that from the movie Her, like that really seamless flowing conversation that you can have with another human. That requires low latency responses. It requires being able to interrupt and requires being able to understand not just text, but, you know, sneezes or coughs or sighs or that type of stuff. And it also requires understanding who is having the conversation. And so Flow uses our best in class ASR. It includes our diarization, which means that it can recognize who it's speaking to. And we built a conversational engine part that enables it to have this very fast, rapid interaction with the large language model and text-to-speech component to enable a truly responsive type of conversational assistant.

24:33So, yeah, we're really excited to be hopefully launching this on Wednesday this week, so in July, and bringing it to some of our early adopters and customers very soon. And the core of, excuse me, the knowledge base and the reasoner is, are you, do you have your own LLM in there? I'm sorry if you just said that, but, or are you making an API call to ChatGPT or Claude or something? Good tech solves problems that you've thought about. Great tech solves problems that you haven't even thought of. What can the commerce platform trusted by millions of merchants do for you? It's time for Shopify, the commerce platform revolutionizing millions of businesses worldwide.

25:28Whether you're a garage entrepreneur or IPO ready, Shopify is the only tool you need to start, run, and grow your business without the struggle. Shopify puts you in control of every sales channel. So whether you're selling satin sheets from Shopify's in-person point of sales system or offering organic olive oil on Shopify's all-in-one e-commerce platform, you're covered. Shopify powers 10 % of all e-commerce in the United States. And Shopify's truly a global force, powering Allbirds, Rothy's, and Brooklyn, and millions of other entrepreneurs of every size across over 170 countries. Plus, Shopify's award-winning help is there to support your success every step of the way.

26:17So sign up for a$1 a month trial period at shopify.com slash IonAI. That's Shopify, S-H-O-P-I-F-Y dot com slash IonAI, that's E-Y-E-O-N-A-I, all lowercase, all run together. Go to Shopify.com slash IonAI to take your business to the next level today. Give them a try. They support us, so let's support them. Yeah, so we've built it to be LLM agnostic. You know, I think what we've seen in the LLM space is rapid development from a lot of different players. And I imagine lots of customers have a preference as to which one they would like to use. And so what we've done is we've brought Llama 3 in into our service so we can self-host it ourselves.

27:15And so that can adhere to any privacy, security, compliance concerns that anybody may have. But you can easily plug that into GPT or Claude or Cohere or Gemini, whichever one you would prefer to use. We don't plan on doing our own LLM. I think there's a lot of resources and other people building bigger and better LLMs. Our domain expertise is really understanding that audio and that speech, and I think that's the opportunity for Speechmatics is that these LLMs are only going to be as good as what they can hear. we want to make sure they can hear really really well yeah uh and i i want to talk more about flow but uh on the research side uh do you uh can you uh handle multiple uh voices simultaneously and parse them.

28:18I had a guy from Google on a year or so ago who was working on bird calls, and they were working on, you know, when you're listening to cacophony of birds in the jungle or deep forest, of, you know, pulling apart that sound and identifying individual streams. And I also am sort of a backyard bird watcher. I have the Cornell Ornithology Lab app. I don't know if you've ever used it. It's really remarkable. And if there are 10 birds, it will list all 10. So it's able to differentiate. So can you do that with human speech? Yeah. I would say there's a spectrum. We definitely are able to understand different voices.

29:20If I got outflow now, it would recognize my voice, it would recognize your voice, it would be able to say who said what. In a busy environment, it would still recognize our voices despite an array of other people around us. Then there's, so you get to the noisy environment and then there's really strong cross talk i don't know maybe in some u.s political situation where there's just lots of people shouting at each other it's hard for a human to recognize uh all of that's going on so that that's still a research project for us to get to in terms of how do you truly pull apart the words that were said by all the different voices at that one time um so we haven't got to that level of true pulling apart, really deep cross talk.

30:09But in a call center, I don't know, with an irate customer where they're talking over the employee that's trying to calm them down, we're able to do that. And so there's a research trajectory on that front. But I would say that we have the best in class diarization for speech recognition. And I would also argue that that is a really critical step in having an assistant that can actually help you in real world scenarios. So if I was walking down the street with talking to ChatGPT and somebody else walked past, it would pick up what they were saying and it would think that that was what I was saying as well.

30:52Whereas with diarization, it knows to listen to only me. You can also then use that to set up things like only respond to me so you know you've got an assistant it's hearing the whole conversation that's going on but it'll only take an action or respond to a question if i say it rather than if anybody else in the room says it actually uh uh you know i wear hearing aids uh and that's one of the challenges of of hearing aids is uh in a in a party or something you just you can't uh differentiate between voices very well. Are there applications like that for hearing augmentation? I think that's a very good application for us to think about for the future.

31:42I know that our founder, Tony Robinson, also wears a hearing aid. And so it's certainly a strong area of focus for him. I think that one of the opportunities here is currently we've built this first incarnation of Flow to be very focused on the mobile device space. I think that's a great place for people to test it, explore it, find what works for them. But I'd be really excited about thinking about putting it into something with multi-directional microphones. So then you can start really getting an extra layer of information out from the diarization, which I'd imagine will work very well in the hearing aid case, where it would listen primarily to people that maybe you're looking at, and maybe tone down some of the stuff that is happening in your periphery.

32:39Yeah. So Flow is being released as an API or as an iPhone or a smartphone app? or how are people going to interact with it? Yeah, so we're very similar to our ASR in that we provide an API such that people can build speech recognition into any product that they're building today. We have a online portal on our website where people can go and play with it and test it themselves, but we don't do a vertically integrated product ourselves. We instead enable others to build that into whatever products they are taking to their customers. And I think that's ultimately what we want to do here is we want to say, we want to enable any product, any technology to be seamlessly integrated with the thing that comes most natural to us, like our voice.

33:39So how do we enable any product to suddenly become keyboard-free, mouse-free, touchpad-free, and suddenly become usable using your voice? I know that a lot of people don't use Siri because it's frustrating sometimes when it doesn't recognize what you say. And so it's kind of surprising that we don't have more products today that are leveraging our voices. We think flow is our inflection point where we get to the point where technology can be used with our voices. We're going to release flow as a conversational AI API so that anybody and everyone can integrate their product and enable their product to be interfaced with through their customers' voices.

Read the full transcript

34:27Now, we're also going to release a sort of app version of it, but that's mostly to sort of showcase the technology, enable people to understand what is now possible, rather than having to like plug and play an API just to be able to test it. Yeah. I think it's a little bit like, you know, OpenAI have released ChatGPT for everyone to use and play with, but ultimately they're selling an API to enterprise customers to enable them to integrate large language models into their products. Yeah, yeah. And does iOS allow the microphone to be taken over by another ASR? So you can build into an app a different ASR.

35:18You still have to use the microphone that's built in or the multiple microphones that are built in. But in an app setting, you're allowed to use different things other than the iOS dictation. Obviously, other people would build on top of what we're offering to provide actions, right? So if Apple were to use our service here, the conversational AI service, we would become the voice interface for Siri. and then Siri would still take all the action calls to tell you what is the weather today and set a timer for five minutes and, you know, all the standard things that you do with Siri. Yeah. Are you talking to Apple?

36:05Oh, well, I think we should be. Hopefully after this podcast, they'll get in touch. Yeah. And since this is, I mean, I do video, but I also do, this also goes out in audio format. Can you give us a demo while we're recording? Yeah. Hopefully my team has finished doing a load balance test. Okay. If not, that's okay. But I'm just getting it open here. So I will just press play on here and we can get started. Hey, Humphrey. can you respond to me and only to me, please? Sorry about that. Yeah, no problem. I'll make sure to respond only to you from now on. How's it going then? That's enough for now.

37:03So what I've done here is I put it into passive mode. And so it's still listening to what we're saying. It's still transcribing what's being said, but it will no longer respond to me. But you saw from that first interaction, like how quick and how seamlessly it responded to my first interaction. Yeah, yeah. And it won't respond if I say, hey, Humphrey, how are you today? It's not going to respond. Well, I think it may have woken up from that. But if I now said it strongly, so, hey, Humphrey, are you listening to me? Yeah, I'm listening. Great. Great. So there are multiple speakers in this conversation we're having.

37:50Please can you only respond to my voice and not to the other person in the conversation? Is that okay? Yeah, no problem. I'll only respond to you from now on. Yeah. So, hey, Humphrey, what's the weather like in New York today?

38:12uh can you could you potentially uh tell humphrey a little bit more about yourself and maybe a bit about this podcast yeah so i'm a former new york times correspondent i got interested in ai after meeting jeff hinton for a story that i wrote and i ended up starting the podcast so that i would have the opportunity to speak to researchers without having to promise that I was going to write an article. Okay, great. So Humphrey tried to interrupt you a couple of times there. I think the cough may have set it off. But if I said, Humphrey, who is it I'm speaking to at the moment? Oh, yeah, I'm Humphrey.

39:00Nice to meet you. And who else is in the conversation with us? there's you and then there's two other speakers let's call them the second speaker and the third speaker that's great and uh let me just pull up something uh one of the speakers uh comes from a particular journal could you tell me which journal that is the second speaker i think is a former New York Times correspondent? So the system can remember what's being said and keep it back and still assigns it to a speaker. So we could also do this where we tell it our names and then it remembers that I'm Trevor and you're Craig rather than speaker one and speaker two.

39:47I think one of the potentially powerful demos that I'd love to show you, if we just have two minutes here, is there was a recent article about McDonald's using IBM to do drive-through recognition and the demo didn't go so well. And so I haven't sort of programmed this in any way for this. I'm just going to prompt it via my voice to be a McDonald's assistant and then we'll see how it does. Is that all right if I try that? Yeah, yeah, sounds great.

40:21Hey, Humphrey. I'd like you to pretend to be a McDonald's drive-thru agent. Is that all right? Oh, yeah, that's all right, mate. I'll do my best to take your order and all that. Welcome to McDonald's. Can I take your order today? Yeah, could I get a Big Mac meal, a large with a Coke? Could I also get six cheeky nugs? And can I also get a bacon McFlurry, please? okay so to confirm you'd like a big mac meal large size with a coke six piece chicken mcnuggets and a baker mcflurry right yeah that'd be that'd be great please do you do you definitely do break bacon mcflurries oh yeah we do the baker mcflurry it's a popular one i'll just put that all together for you would you like to try our mobile payment and pickup or would you like to pay at the window I'll pay at the window.

41:16Thanks very much. No problem at all. Your total comes out to be, let me just check, asterisk, pause asterisk, of£11.99. I'll have everything ready for you at the next window. Please drive up and we'll get you all fixed up. Thanks very much, I'm free. That's interesting. So a couple of questions. First, is it only responding with the wake word? So you can set it up in multiple different ways. You can set it up to only respond when you say the wake word Humphrey, which is the persona that we've currently given it. You can give it any persona you'd like. or you can set it to a much more lively mode where it's always on, always responding.

42:03And or you can even set it to not have a wake word, but try and work out when is the right time to engage in the conversation. This to me is that that's where the power of large language models really comes in and that you suddenly have this generalizable tool that can try and work out when is the right time to interact in the conversation. Yeah, well, on that point, do you know the company Tenex? T-E-N-Y-X? I had that one on the podcast, but they're working actually with fast food places on this kind of solution. And one of the problems that they have focused on is endpoint prediction. Like when does the model know that you finish speaking and that you're not just pausing?

42:50Is that something that you guys have spent time on? Yeah. So when we did the first version of this, we tried to look into the endpoint prediction in the ASR model. And we already do that fairly well for including punctuation in our ASR. But what we found was that when you get to the conversational AI element, it creates this unnatural latency when you finish speaking. And so what we actually found was that instead, if you're finding the right way to put the information from the ASR into the LLM as quickly as possible, and then enabling a few different approaches to understanding when is the right time to respond, that gives you this much more seamless interaction.

43:39and then the other really critical point is interruption right so with chat gbt if you're speaking to it when it's speaking back you can't interrupt it you have to touch the screen to interrupt it sorry there's a there's a i don't know if you can hear the ice cream truck in the background uh whereas with ours what we built in is this ability to interrupt it with your voice So if it does accidentally start before you finish, if you're a slow talker or if you pause for intonation, it does try and respond. If you carry on speaking, it will just naturally stop. It will continue to listen to the whole conversation and it will merge that whole section together.

44:21So I'll understand that it's not a completely different request between the break. Yeah. And this is kind of funny because I, as I mentioned, I've been, I built something that's, it's also model agnostic as kind of a, you know, I drive a lot and, and I'm always curious about what's around me. And so I wanted something that, that would give me the history of the places I'm driving past. And it's got a wake word. And I mean, it's, it's just for me right now, but, um, uh, and it's right now it's using GPT-4-0, you know, which hallucinates, but for my purposes, it's, you know, when it's hallucinating, you know, if, and it's, it's pretty good.

45:12But would this do this if you asked Humphrey, what's the history of Chappaqua, New York, for example? Would it, in its knowledge base, presuming that LAMA2's training data includes that, would it find that and respond? Yeah, so I think some of the sort of most fascinating examples of SpeechMax employees trying it have been exactly the scenario you've been talking about where people have been driving in the car, they're bored of the radio, they just want to have a play around and they turn it on and they end up having a conversation about something that they're going past at the moment. So, yeah, Llama 3 appears to be pretty well tapped up to the history of the Cambridge area, which is where we're based, Cambridge and London, and does a pretty good job of that.

46:07Obviously, it does have the same issues as all large language models in that sometimes if it doesn't know it, it will just make something up. But I've also had some fairly in-depth conversations with it about, I don't know, Monte Carlo methods and hidden Markov models and recurrent neural networks. And it does a pretty good job. There's clearly a lot of text about that out on the Internet there at the moment. And the great thing is, is that the voice it gives you back makes it feel like an engaging experience. It makes it feel like you're not just reading a wall of generated text, but instead you're having this deeper conversation and you're actually learning through engagement and interaction rather than, yeah, by having to scroll through a lot of bullet points that come out of you on the screen.

47:01Yeah, yeah. And so this is being released as API to be built into other products, but you are releasing this, the app that you were just demonstrating. Yeah, so in two days' time, we're having our launch event. And so we'll be doing a demo up on stage in front of an audience of, I don't know, I think it's 50, 60 people. I think it's also going to join our panel. And so if that goes well, expect Flo to join a few other panels as well. Maybe in the future, it could come on this podcast and you can do a pod with them. and yeah, hopefully then launching live. If people want to find out more, I'd highly recommend them to go to speechmatics.com slash flow.

47:51That there, you can sign up to our wait list and we can give you comms and information about how to try it out more. But we're very excited to think about how to get this into hands with more people so that even more interesting use cases than I can dream of can come up and come to us. Yeah. And I'm just curious, why not release it as a B2C app? Or I mean, you are in effect, but why not promote it as a B2C app? Yeah. So I think there's a lot of interesting examples of how to like exploring large language models through this interaction. But I think what people want the most is a problem of theirs, a use case of theirs to be solved.

48:36And so there's an awful lot of incredible products already out there in the market that I think could be massively improved if we were able to interact with them through our voice. I think we've probably reached peak chatbot-based apps on our phones. We've all probably got a bit bored of having the same conversation with LLMs. I think the most powerful use cases will come through when we integrate these with products that we use every day, day in, day out, but we no longer have to get out, type something in or touch a screen, work our way through a UI, but instead can just utilize natural language to give our technology a command and it understands us, recognizes the context and takes an action based on what we said.

49:24Yeah. Do you think that, you know, Siri working with OpenAI, I hope eventually with you, that it will become sort of the killer iPhone app that will take actions and answer questions about history and everything like that? yeah i think i think the main limitation of siri like aside from not being able to understand every voice sounding a bit robotic was that it just how structured it was and so it always felt limited in the way you could leverage it to take actions and so that's why everyone just uses it to set a timer for while their chickens right right um because that's the useful use case i think large language models give you the ability to understand more complex natural language and then understand what actions need to be taken there and so i think siri plus gpt is is definitely a massive step forward but i would still sort of say that you know garbage in garbage out you really need to understand what is being said who is it being said by and how is it being said to really be able to get to that killer app which is why i truly believe that you know speech is a core part of any future agi stack yeah and on the various languages what's the most obscure language that you've uh trained it on so far well i don't know um there's uh there's there's a few i guess um I guess we're going into Tagalog at the moment, I think, is one of the ones we're exploring.

51:19Yeah, you've put me on the spot here. I can't think of which one is the most obscure. But I'd probably be offending somebody if I said too many times. There's millions of speakers that speak that language. That's right. I guess obscure is the wrong word. Yeah, but you do Icelandic, Latvian, Burmese. I don't know what else is out there. Yeah, so we have just finished covering off all the languages that are required for the EU. So we've just finished Maltese and an improvement on Irish, which means that we're covered for some EU regulation that's coming in very soon for that. We've recently released Hebrew and Persian, and we've just done a big uplift on Arabic.

52:20One of the things I would say with Arabic is obviously there's a huge range of different languages within Arabic. But what we've been able to do because of our combination of self-supervised with the acoustic model, with the label data, is we built a global Arabic model. And so it can understand all the different languages and dialects within Arabic rather than having to choose which one you want to focus in on. And so, yeah, we're particularly excited to be releasing those and watch this space for more. more um you can always go on uh to the speech max website and check out our list of 50 languages if there's uh something in there you want to explore yeah and that's uh speechmatics.com is that right that's right on the on the multi uh sort of uh pulling apart uh uh voices in a in a crowd Where are you on that?

53:20Because that's, I mean, you mentioned national security applications. That seems like it would be a huge benefit. it? Yeah, so we offer diarization, which is the term for understanding different people speaking on top of our ASR as a standard feature for anyone to use. The scientific method we use for doing that is a form of clustering technique where we sample and we try and cluster on some Disney plot a different distribution of the speakers. And so if you have two different speakers, like in this conversation, then obviously you get quite a clear separation as long as we don't sound too similar.

54:11And so it's then generally quite accurate in terms of recognizing which speaker is the one that's been talking. As you increase the number of speakers, you get some more and more overlay in the cluster space. So we go up to, you can do as many as you like, but I think if you went more than 20 speakers, you start to see quite significant degradation in terms of recognizing which speaker it is. And so, yeah, it's a big part of our next research project as well, alongside the multi-scale, is that we're also exploring how to get ever-improving diarization results for more and more speakers. Yeah. And how big is the team at Speechmatics?

55:01Yeah. So the company is about 120 people. There's about 60 people on the engineering side. And of them, there's probably 20 or so machine learning researchers and engineers. So as a company, we're quite sort of heavy on the research side. I would consider us a bit of a research lab, a bit of an applied research lab where we're always doing fundamental research, exploring other fundamental research and thinking about how do we bring this into the speech domain. And that's enabled us to provide multiple revolutions in ASR from the early days, where everybody was using Kaldi all the way through to using Transformers today.

55:46And I expect us to continue to explore and invest in ways that we bring some of the latest and greatest AI breakthroughs into the speech domain. That's it for this episode. I want to thank Trevor for his time. If you want to learn more about the conversation today, you can find a transcript on our website, IonAI, that's E-Y-E hyphen O-N dot A-I. You can check out Speechmatics at Speechmatics dot com. And remember, the singularity may not be near, but AI is changing our world. So pay attention.

From the publisher

In this episode of the Eye on AI podcast, we explore the forefront of voice-powered AI technology with Trevor Back, Chief Product Officer at Speechmatics. Discover how Speechmatics is pushing the boundaries of speech recognition and conversational AI with their latest innovation, Flow.

 

Trevor shares his journey from a background in computational astrophysics to becoming a key figure in AI at DeepMind and now Speechmatics. He delves into the development and potential of Flow, a groundbreaking tool combining automatic speech recognition (ASR), large language models (LLMs), and text-to-speech synthesis, aimed at creating seamless and responsive voice interactions.

 

We explore the wide-ranging applications of Speechmatics' technology across industries, including media, call centers, and education. Trevor discusses the challenges of achieving high accuracy in speech recognition, especially in diverse and noisy environments, and how Speechmatics addresses these challenges with their unique approach to training models.

 

Listen in as we uncover the intricacies of handling multiple languages, improving diarization, and the future goals of understanding complex audio cues like emotion and sarcasm. Learn about the company's vision for integrating voice technology into everyday products, making technology more accessible and user-friendly.

 

Don't miss this insightful conversation on the future of voice technology, AI in business, and its role in the evolving landscape of AI. Like, subscribe, and hit the notification bell for more expert discussions on cutting-edge advancements in AI.

 

 

This episode is sponsored by Shopify.

Shopify is a commerce platform that allows anyone to set up an online store and sell their products. Whether you're selling online, on social media, or in person, Shopify has you covered on every base. With Shopify you can sell physical and digital products. You can sell services, memberships, ticketed events, rentals and even classes and lessons.

Sign up for a $1 per month trial period at http://shopify.com/eyeonai

 

 

Checkout Speechmatics, the most accurate AI speech technology - with AI transcription & real-time translation components.: https://www.speechmatics.com/

 

 

Stay Updated:

Craig Smith Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI

 

 

(00:00) Introduction and Background

(01:49) Trevor Back's Journey into AI

(04:02) DeepMind and Early AI Applications

(07:30) Speechmatics' Mission and Focus

(12:06) Key Applications of Speechmatics Technology

(14:25) Achieving High Accuracy and Low Latency

(17:52) Language Coverage and Challenges

(21:27) Future of Voice Technology and AGI

(24:52) Integrating Large Language Models

(27:31) Handling Multiple Voices and Diarization

(29:32) Real-world Applications and Challenges

(35:20) Demonstration of Flow and Capabilities

(41:14) Endpoint Prediction and Interruption

(43:53) Real-time Interactions and Future Prospects

(45:34) Launch Event and Future Plans

(50:13) New Language Releases and Compliance

More from Eye On A.I.

All 266 episodes
#200 Trevor Back: How Speechmatics is Shaping the Future of Conversational AIEye On A.I. · 56 min
Listen in VO