In short
Eye On A.I. Podcast Episode Notes
Episode Title
#166 Itamar Arel: Is Voice AI the Future of Customer Service?
Host
Craig S. Smith
---
Episode Overview In this episode, Craig Smith interviews Itamar Arel, CEO of Tenyx, a company focused on creating advanced voice-based conversational agents using neuroscience-inspired AI technology. The discussion centers around the evolution of voice AI in customer service, the challenges faced, and how Tenyx aims to reshape the customer service landscape through innovative technology.
---
Key Discussion Points
- Itamar Arel's Background
- Transition from Academia to Entrepreneurship:
- PhD in Computer Engineering; postdoc at Stanford.
- Faculty position at the University of Tennessee, focusing on machine learning and AI.
- Shifted to industry to pursue practical applications of his research.
- Journey with Tenyx:
- Founded Tenyx after experiences at Apprente, which was acquired by McDonald's.
- Arel’s vision is to develop robust voice AI agents for automating customer service tasks.
- Current Trends in Voice AI
- Evolution of Technology:
- Advances in voice AI and large language models (LLMs) have changed the landscape of conversational agents.
- Arel emphasizes the ability of LLMs to understand natural language better than earlier technologies.
- Challenges in Voice AI Development:
- Understanding human speech nuances and maintaining a natural flow in conversations.
- The need for robust Natural Language Understanding (NLU) to accurately interpret user requests.
- Tenyx's Approach to Voice AI
- Operational Framework:
- Arel explains how Tenyx creates human-like voice interactions by customizing open-source models to minimize latency—key for real-time voice interactions.
- Eliminating Hallucinations:
- Tenyx employs a multi-guardrail approach to prevent erroneous outputs, ensuring adherence to business logic and accurate responses.
- Market Position and Differentiation
- Navigating a Crowded Market:
- Tenyx differentiates itself by focusing on voice interactions, which are less saturated compared to text-based chatbots.
- The company leverages past experiences to build trust and deliver effective solutions.
- Four Pillars of Effective Voice AI:
- Robust understanding of user speech.
- Accurate response generation.
- Fine-tuning capabilities to mitigate issues.
- Optimizing voice dynamics for natural conversations.
- Future Insights and Predictions
- Five-Year Outlook:
- Arel expresses optimism about ongoing advancements in AI technology, including improvements in speech synthesis and multimodal models.
- Continuous research efforts aim to develop LLMs that can optimize actions over extended horizons rather than just predicting the next token.
---
Key Takeaways
- Itamar Arel’s journey illustrates the transition from academia to impactful tech entrepreneurship.
- The future of customer service is likely to be heavily influenced by voice AI, with companies like Tenyx pushing boundaries.
- Understanding human nuances in speech is crucial for creating effective conversational agents.
- Continuous evolution in AI capabilities presents both challenges and opportunities for innovation in customer engagement.
- A rigorous evaluation process is essential for maintaining quality and reliability in AI-driven solutions.
---
Conclusion The episode highlights the transformative potential of voice AI in customer service and the innovative approaches companies like Tenyx are taking to harness this technology. As AI continues to develop, its impact on industries will only grow, making it imperative for businesses to adapt and innovate.
---
Additional Information
- Sponsor: Shopify, a platform enabling entrepreneurs to set up online stores and sell products.
- Stay Updated: Follow Craig Smith and Eye on AI on Twitter for more discussions on AI advancements.
Listen to the full episode for deeper insights into voice AI and its future in customer service!
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00In the brain, as we humans or mammals really learn new tasks, we tend to invoke parts of the brain so as to learn this new task while other parts of the brain remain dormant or inactive. And that makes a lot of sense. That's what allows us to learn how to ride a bike without forgetting how to walk. What it involves is sort of processing the information, looking at the geometry essentially that's being represented inside the network. Networks is this deeply layered billions of parameters. And so if you have a way of looking at which parts get activated as that new information comes in, and you have statistics or information about which parts were called upon when the original dataset was trained on.
0:37Then you have sort of a chance at figuring out which neurons, which weights should be used in fine tuning and sort of mask off or deactivate the rest. This is super high dimensional representation, so it's hard to visualize. The underlying thing I would say is that these networks are really linearly partitioning the space. If you're able to understand how that happens and analyze these, what are called the fine transforms, these patches of partitions, then again, you have a chance of figuring out which partitions should remain as is and kind of not move those boundaries, whereas in contrast to the ones that should be adjusted so as to learn this new information.
1:15Hi, I'm Craig Smith, and this is Eye on AI. My guest today is Itmar Arell, a former academic and now a tech entrepreneur who's been at the forefront of voice AI customer service solution. Itmar shares his journey from academia to the tech industry, detailing his work at Tenyx, T-E-N-Y-X, a company specializing in human-like voice AI experiences for customer service. Itmar delves into the intricacies of developing robust voice AI solutions, addressing the challenges of understanding and responding to human speech and the nuances of voice dynamics. I hope you find the conversation as enlightening as I did.
2:02Hi. Good tech solves problems that you've thought about. Great tech solves problems that you haven't even thought of. What can the commerce platform trusted by millions of merchants do for you? It's time for Shopify, the commerce platform revolutionizing millions of businesses worldwide. Whether you're a garage entrepreneur or IPO ready, Shopify is the only tool you need to start, run, and grow your business without the struggle. Shopify puts you in control of every sales channel. So whether you're selling satin sheets from Shopify's in-person point of sale system or offering organic olive oil on Shopify's all-in-one e-commerce platform, you're covered.
2:46Shopify powers 10 % of all e-commerce in the United States, and Shopify's truly a global force, powering Allbirds, Rothy's, and Brooklyn, and millions of other entrepreneurs of every size across over 170 countries. Plus, Shopify's award-winning help is there to support your success every step of the way. Sign up for a$1 per month trial period at shopify.com slash IonAI. Give them a try. They support us, so let's support them. Itamar, introduce yourself, some of your background and how you got to Tenex and what Tenex does and what you guys are working on. Yeah, well, first, thanks for having me.
3:37I guess a little bit about my background. I still consider myself a recovering academic on most days. After getting a PhD in computer engineering, I, in a short postdoc at Stanford, took a faculty position at the University of Tennessee in the electrical engineering and computer science department, where I ended up doing 10 years of research in machine learning and AI, an area specifically pertaining to reinforcement learning and what later became the field of deep learning. Mind you, this is just before the 2012 deep learning revolution, so a very different time, much smaller community. And then after that, eventually took a couple of years of sabbatical back at Stanford, had a courtesy appointment, visiting professorship at the AI lab and worked with Silvio Savarese in a fairly large DARPA project we had.
4:28Silvio, of course, at the time was a CS professor and since took the role of chief scientist at Salesforce and enjoyed all that. But at some point realized that what I was really passionate about is taking research outcomes and maybe pushing them a little further and seeing what products and services they can empower to really impact a lot of people. And when you kind of reach that recognition, you realize that academia might not be the optimal setting for you. So did the uncommon, you know, gave up tenure and left academia, moved to the Bay Area. I was an EIR, an entrepreneur in residence for a while at Ame Cloud Ventures, which is Jerry Adams Fund, the guy who started Yahoo, of course, and then started a company called Apprente in 2017, where the vision very early on became building these robust voice AI agents to automate the order taking process at drive-thrus.
5:22So think Starbucks, McDonald's, all these chains that we know and love. The idea was to fully automate that drive-thru ordering process. Interestingly enough, across what's called the quick service restaurant chain world, over 70 % of the revenue comes from the drive-through, which is a little mind-blowing for most people that hear that first. But it was an interesting use case. Of course, this was pre-LLM, pre a lot of the revolution or the advancements of the last two, three years. We ended up working with a fair amount of these well-known chains. And in October of 2019, we were acquired by McDonald's Corporation.
5:59The CEO at the time believed that McDonald's views itself rightfully as an industry leader. So they believed in this solution, this project, and they were interested in the team and the tech, I suppose. But also they wanted to establish a Silicon Valley center of excellence around intelligent automation, AI. That was kind of the drive, the vision of the CEO at the time. And so that was sort of the pitch to us. How do you feel about being the founding team around which we'll build this center? And it did unfold that way. We were about 20 at the acquisition and grew to about 100, about a third PhDs, very focused on this voice platform solution.
6:35A really interesting and fascinating journey. This solution is now deployed in hundreds of stores. I think the plan is still to deploy this across 14 ,000 stores in the U.S. and then go global. To maybe complete that chapter, two, three years later, under the leadership of the new CEO, there was a sense that maybe they're still very bullish about this initiative, but felt like maybe the right house for that organization, which again was about 100 people at the time, was at a tech giant rather than McDonald's. So at the end of 2021, we were reacquired, the whole organization joined IBM, became part of IBM Watson, and it's still part of IBM Watson.
7:17And now IBM, of course, offers this to all these chains. So spent some time there, but at that point uh you know if you're hit with the startup bug uh then you you kind of miss it and some of us felt like we missed being in a small setting a small group that wants to do something new and change the world so uh that sort of led me or led us to tenix about almost two years ago are you now competing with ibm and mcdonald's or or is your solution different from what you were doing there Yeah, totally not competing, totally different solutions. I mentioned that really the last two, three years have been transformational in the field, right?
7:55So LLMs weren't around much of the advances around ASR, the speech recognition side, TTS, as well as really every part of the pipeline is new and unrelated to that. I just showed you how fast the field is moving, of course, because it's only been a few years. So, yeah, we don't operate in the restaurant space, and it's based on totally new technology. Yeah. Before we go on to the new technology, what happened to the order-taking with a conversational agent? I mean, I've never run across that in a drive-through. Yeah. So McDonald's is based in Chicago, and so deployment started from the Midwest.
8:37That's sort of the epicenter, and then it's expanding from there. It is deployed in hundreds of stores. I believe the plan, I think it's still not 100 % sure, but the plan by early 2025 to be across 14 ,000 stores, which is the McDonald's stores in the US, and then explore going global. McDonald's has almost 40 ,000 stores worldwide. As I mentioned, good amount of the revenue of the order volume actually does come from the drive-thru, so it makes perfect sense from a business standpoint. Yeah, it's live and kicking. So how does 10X's technology differ from what you were doing before? Yeah, so as I mentioned, there's been a lot that's happened in the ML world since then.
9:22Of course, the elephant in the room is the introduction of large language models. And what that really allows to do in a much easier way than it was before is understand, introduce robustness to the NLU part of the pipeline, the natural language understanding. I always give this example. People can often go to the drive through and say something like, can I get a number two? Yeah, number two, a Diet Coke. So you and I know that was not two orders. That was just a completion of thought. That was one order. But in order back in the day, back in the day, quote unquote, 2020 and before, to try to capture all these million different ways people have of asking or ordering their food or asking for information, you really needed to either templatize the various ways in which they can ask it or simulate it.
10:09It was quite challenging, in fact, to get the long tail of, again, ways in which people communicate. LLMs really changed all that in the sense that you can actually try it on ChatGPT. If you say, customer said, quote, can I get a number two dot, dot, dot? Yeah, number two, a diet coat. What did the customer ask for? And in all likelihood, you're going to get one number two, a diet coat, right? And so first and foremost, I think NLU has contributed to making the understanding, the robustness and the understanding much more, much easier and quicker. And we did eventually get to automate the vast majority of orders at McDonald's.
10:49But frankly, it just involved a lot of manual work, a lot of laborious designing when now LLMs can really help out with that. Yeah. So TenX, when did you start at TenX again? So TenX started in early 2022, almost two years ago. So you started with generative AI, with LLMs, is that right? Right. Exactly. We knew about it slightly before it was fashionable and felt like that was one of the big things that can really change the field in a sense. So TenX, we offer our customers what we feel is truly human-like voice AI customer experience. And customers tend to be mid-sized to larger customers that are interested in either completely automating the voice, their customer service or voice-based customer service or just the vast majority of it.
11:45And basically, there's really a bunch of things you have to do that ties into your previous question to do that well. I mean, I'm sure we've all had the experience of calling a hotel reservation or airline, and you're usually greeted. If it is an automated system, you're greeted by a system, a conventional system called IPRs. And these voice response systems tend to be very brittle. It's the ones that tell you, you know, press one for that, press two for that. I think you said representative. And what winds up happening most of the time, I'm sure you and I both experience this, is 90 % of the people actually press 000 and just ask to talk to a person.
12:23And the strong sense is that now in 2023, soon 2024, with the introduction and particularly with all these new technologies, we can finally build systems that can feel natural and flowing and robustly understand us and provide accurate responses to the point where it's just a much better experience and you don't press 00 90 % of the time. Maybe 90 % of the time you're serviced well by these systems. And that's sort of the vision. Yeah. A couple of questions. Are you guys, do you have your own model or are you pinging a commercial model through an API? Yeah, so we have to own our own models, partly because of latency issues in voice.
13:07Voice conversation is very sensitive to latency. You can't have like five seconds between turns. It really needs to be below second and a half, ideally even second to feel natural. And so we end up taking open source, cutting edge open source models and customizing them in many ways to make them work well for us. So, yeah, we own our models. Yeah. And the other obvious issue is hallucinations keeping the model on point. Do you guys use retrieval augmented generation or some strategy, vector database strategy? Yeah, so absolutely. I think the conventional wisdom is that to make the most of LLMs, you do need RAG, but you also need fine tuning to make this thing work really well in any specific domain.
14:01I think that's one of the relaxing aspects of the problem we're tackling is that on the enterprise side, almost every company, the conversations that people call and talk about tend to be within this restricted domain. 99 % of the people that call an airline call to either book a flight or change a flight or cancel and so forth. So that makes it slightly easier in contrast to consumer facing products like Siri or Alexa. that is just super open-ended, right? You can ask about music, you can ask about information of the web or the weather in San Francisco tomorrow. So you touched on hallucinations.
14:34In our world, you're absolutely right. There's just almost zero tolerance to a solution that doesn't adhere to business logic and rules. You can't say things that are not grounded in reality. You certainly can't say things that are offensive or otherwise inaccurate. And so we developed an approach that has multiple guardrails to mitigate that problem. Some of them rely on traditional rule-based schemes, and others are deeply rooted in machine learning techniques. And the spirit of things is to try and characterize a normalcy profile, a distribution, if you will, of what conversations in any particular domain look like so that you can then use sort of anomaly detection style techniques to detect whether the conversation somehow veered off.
15:25It is now off topic or something is not within the scope of what it should be. And so the multi-guardrail approach, in our humble opinion, is the way to reach, you're never going to reach zero, but you're going to get arbitrarily close to 100 % accurate or grounded conversation that it does not hallucinate. But this is an ongoing challenge and a problem for just the community at large. Yeah. And do you specialize in one or two verticals or is this a horizontal solution? Yeah. So we started with travel and hospitality. So hotel reservations and things of that nature, and then expanded to real estate, insurance, fintech, the nice thing we think about the particular approach we've taken is that it can fairly easily be applied across multiple domains and that's exciting.
16:20Part of it's powered of course by this architecture that incorporates LLMs, other has to do with the other pieces of the puzzle. I'll say one thing about fine tuning because I think you mentioned that. Part of the work we've done and we just had a piece in VentureBeat about that is fine tuning is as it sounds, a way of taking some amount of information within a restricted domain, say hotel reservations. So you take a bunch of conversations of real world conversations between an agent and a customer. And what you'd like to do is take that so as to polish, to customize an off the shelf sort of foundational LLM to be better in that domain.
17:01And fine-tuning just involves changing the weights and in effect sort of continuing the learning so as to change the weight so that, again, it's a little better in that domain. It captures the right nomenclature. It responds in a prototypical way. The challenge we noted almost a year and a half ago when we started playing with these things is that it's not as easy or as simple as it sounds. When you change the weights in the process of fine tuning, you oftentimes end up distorting things that were there before. So you suddenly forget, forget knowledge, forget some reasoning. And that's called, that relates to a well-known problem machine learning called catastrophic forgetting which has been around for 30 40 years how do you continue to train while making sure you retain what was there before and so we have a research team that sort of took that on because we realized that was a key challenge something we had to solve and came up with um with a novel scheme that was that's able to do that to to allow for fine-tuning while making strict guarantees that the knowledge and the reasoning and more importantly, the RLHF protection that was there before remains.
18:08Because one of the other things we saw is that oftentimes when you fine-tune, even on benign in-domain conversations, suddenly the protection that was there introduced by RLHF can be removed or cracked to the point where either harmful or biased comments can be produced. And of course, that's highly undesirable. Can you talk about that novel approach? Have you published on that? water because that's certainly a hot topic right it is rightfully a hot topic um you know we haven't published it yet we are offering it as a beta service uh which people can sign up on our website um but to share a little bit about the sort of the spirit of of the approach we've taken it's a bit of a neuroscience inspired approach in the sense that there's um fair amount of consensus that in the brain as we humans or mammals really learn new tasks we tend to invoke parts of the brain so as to learn this new task while other parts of the brain remain dormant or inactive.
19:07And that makes a lot of sense. That's what allows us to learn how to ride a bike without forgetting how to walk. And that's not been the traditional approach to continue learning or fine tuning. And we were sort of inspired by that in the sense that when fine tuning, we have an approach that allows the network to look at the data coming in, figure out which parts, which neurons, which weights should be included should be invoked you know tune those so as to learn as most as much as possible from this data in this in a specific domain while kind of keeping everything else as is and that allows you to fine-tune while mitigate forgetting of the get knowledge reasoning and the rlhf stuff i'm sure there's there's more than one approach to do it uh but we've uh we've come up with a framework that we think um is is very scalable and that's been pretty exciting yeah Can you talk about the framework?
19:58I mean, what it involves? Is it offloading memory to a knowledge base? So not quite. What it involves is sort of processing the information, the fine-tuned, the information to be fine-tuned on, and looking at the activations, looking at the geometry, essentially, that's being represented inside the network. Network is this deeply layered, you know, now to the billions of parameters. And so if you have a way of looking at which parts get sort of activated as that information, this new information comes in, and you have statistics or information about which parts were called upon when the original data set was trained on, then you have sort of a chance at sort of figuring out which neurons, which weights should be used as in fine tuning and sort of mask off or deactivate the rest so as to retain what was there before.
20:53And in a sense, that's a deeply mathematical approach to look at the geometry of what was represented before and what's being called upon or represented as this new data comes in. Oh, that's fascinating. So it's a mathematical, you don't have a visualization of the network. No, that's a challenge. This is super high dimensional space representation, so it's hard to visualize. But the underlying thing I would say is that these networks are really linearly partitioning the space, right? So if you're able to understand how that happens and analyze these, what are called the fine transforms, these patches of partitions, then again, you have a chance of figuring out which partitions should remain as is and kind of not move those boundaries.
21:43whereas in contrast to the ones that should be adjusted so as to learn this new information. And that's maybe a little technical, but that's sort of the underline. Yeah, yeah, that's fascinating. And the fine-tuning, the data that you use for the fine-tuning is, yeah, where are you getting the data sets? Is that stuff that you, are there data sets or customer interaction data sets out there that you can license or that are open source? Or do you build up your own or do you use synthetic data? All of the above. I think the best approach is to use real data, even a little bit of real data, and have a way of taking that real data and using a high-performing LLM like GPT-4 to generate synthetic data that covers that space, that distribution, right?
22:40So some way of taking real data, which you always have less than you'd like, and having that be the core to generate additional synthetic data and then train on all, train and evaluate on all that, we think that's probably the best approach at the time, and that's generally the approach we've been taking. Yeah. And that's the training data is text, not voice, right? I mean, the voice to text and the text to voice on either end are not part of the inference. Right. We partner with other companies to do the speech to text, the ASR and the text to speech. There's a lot of companies that specialize in that.
23:23I guess the one thing I would say is in additional to the transcription, to the text, which of course is the core of what we're processing, we do extract non-fonatic, non-textual information that oftentimes gives you some just additional cues as to what the customer was saying. And in particular, it pertains to The other thing I think that's key to getting this kind of system to work well is to handle just the dynamics of the voice exchange well. We speak very different than we write or we text. We tend to use poor grammar and broken English and ums and ems and repetitions. But somehow, you know, a six or seven-year-old can handle all that almost naturally.
24:03And traditionally, it's been challenging to build machines that do that well, right? And so we really focused on that because we think that's key to making people, to engaging and having a flowing conversation as opposed to having people press 000 and ask to talk to a person. Some of it has to do, for example, you sort of alluded to that, to detect when is the person talking, kind of finished saying what he or she wants to say versus when are they just mid-thought, right? So somebody can say, imagine a hotel reservation use case where the customer says, well, I'll be checking in on, you know, then he or she kind of checks their calendar and three seconds later, so on the 15th and departing on the 18th.
24:43This is things we say every day, right? But understanding or detecting that's called endpointing prediction in the world of voice processing. That becomes critical because if you somehow start processing half that statement, you're likely to produce garbage and the whole conversation can go the wrong way. Other times when somebody just asks, well, what's your checkout date or checkout time? You do want to detect that really quickly and say, checkout is at noon, right? So those kind of things, dealing with situations when one side speaks over another, we humans, again, have this tendency to figure out how to resolve that, say, oh, I'm sorry, go ahead, or understand whether the other side finished hearing what I was saying or do I need to repeat myself?
25:27All these nuances of the voice dynamics are really, really critical to get these things to not be brittle and really be engaging and flowing. And that's a big part of what we're building. Hmm. The market of, I mean, since GPT-4 certainly has kind of filled up with all kinds of solutions, and I would guess 10X's market is as crowded as any of them. How do you differentiate yourselves? How do you carve out market share? Is it largely on the strength of your past relationships? Is it through marketing? I'm just curious for a solution like this because, I mean, I'm in a very different and much less sophisticated activity, but I get pitched daily.
26:27There must be at least five, oftentimes more, pitches from LLM-powered marketing companies of one sort or another. We can grow your podcast to 10x. And it's just, you know, and I'm sure each one has its merits, but there are just so many. So how do you handle that? Yeah, you're right. I mean, even in the customer service automation space, there are a lot of companies. The majority of them are actually in the text realm, like chatbots for e-commerce and like that sort of the natural immediate use case. voice, as I mentioned, voice is more nuanced. Voice is challenging. We speak very different than we write and getting to manage voice conversations as well is a non-trivial task.
27:22Maybe that's part of the reason why that subspace like voice AI customer service automation is far less crowded. I'm sure it will get more crowded in the future. And to your point, we do feel like we bring five years of experience in building voice AI solutions in the real world with our McDonald's adventures. And so a lot of those learnings do crossover and particularly around making the conversation feel natural and recovering well. I guess I would say there's four things you got to get right to build an efficient and engaging voice AI solution. One is robustly understanding what the customer said, right?
28:01And that has to do with everything we talked about the broken English and the poor grammar and so forth. The other, of course, is producing accurate responses, right? You do want to respond. When you respond, you want to respond accurately. And that has to do with incorporating RAG and fine tuning. And I would even say fine tuning in the context of RAG, right? RAG involves, or retrieval augmented generation involves having a semantic sort of embedding space that allows you to then search a knowledge base and dig up the right pieces so as to respond to the customer. but that relies on that embedding space being accurate right and differentiating and so there's there's work and so we certainly do that in fine-tuning even the rag piece so that we really make sure we we capture the right representation dig up the right pieces of out of the knowledge base to to respond so that's sort of that was number two three as i mentioned that you you really need fine-tuning end-to-end to do well in any restricted domain and to mitigate things like hallucinations and adhere to the business logic and rules, which is critical for any enterprise.
29:06And finally, everything having to do with optimizing the speech dynamics. So barging and endpointing and dealing with acoustic distortions, all of that is non-trivial. And luckily, yeah, specifically the last two, three years have been transformational and offering these new pieces of the puzzle on the technology side that we strongly feel now you can build the solutions that we've been promised for 20 years that that are robust engaging flowing and create the right customer experience we're sort of obsessed with that i think it all starts and ends with with with really delivering a customer experience that feels like you're talking to a person uh and you were the different uh industries that you're uh that you're looking at do you fine-tune the same model uh so let's say you're in hospitality you fine-tune the model for that industry.
30:01You have a customer in that industry. You take all of their data documentation, if it's restaurant menus and that sort of thing, put it into a vector database and use retrieval augmented generation for the output. But, and then when you move on to, I don't know, another industry, any other industry, real estate, then do you go back and fine tune the same model, you know, being careful not to catastrophically forget the previous fine-tuning and then plug in a new vector database for whatever customer there is in that industry? Or do you have to build, create a new instance of the model and fine-tune it for each industry?
31:09I mean, can you fine-tune? Is it one model that keeps getting fine-tuned And as you add industries, or do you have a row of models, one for hospitality, one for car repair, one for whatever? Yeah, so it's more of the latter. Every industry, we found that you could sort of fine-tune for that industry. And the small differences between, say, Hyatt and Hilton and Marriott usually don't merit complete fine-tuning. But across industries, yes, you do need to specialize. Each has its own terminology and nomenclature and just dynamics of how typical conversations unfold. And so you do need to do that per industry fine tuning.
Read the full transcript
31:54Having said that, one of the things that rightfully any customer is very sensitive to is contamination. I think you alluded to that contamination of their private, let alone PII information. And so we take extra steps to make sure that anything that is unique or private to any customer is never used across even customers in that same industry. But conceptually, yeah, every industry would have its own sort of fine-tuned macro. And you talked about fine-tuning the vector database or the RAG system or mechanism or whatever. I haven't heard anyone talk about fine-tuning on that side. What does that involve?
32:39So it involves the same ingredients, if you will. You want to have typical or prototypical conversations in that specific domain. So that usually means you take whatever actual real world conversations you have and use high performance LLMs to enrich that with synthetic, to augment that with synthetic data. And ideally, you want to take all that and use your embedding model, if it's a bird or similar models, and essentially fine tune, like continue to train that model on that data. So it's really kind of a polishing. It's not that you're training everything, this whole embedding from scratch. You want to leverage everything that that model knows about general language and semantics and so forth.
33:26But you want to refine it such that that vector database becomes more accurate, more differentiating for that particular domain. It's not always needed, but a lot of times it is. And we found that to be beneficial. Yeah. How do you evaluate these models once they've been fine-tuned and you've got your rag working? What's the process for evaluation? Evaluation. Because you can't just then turn it on and try it out, you know, even with a dozen people in the office and, oh, yeah, it worked great. I mean, if you're going to put it in production, you need some robust evaluation to make sure that it's catching the long tail or you know where in the long tail it drops off, that sort of thing.
34:28Absolutely. And the other reason why you don't want to just test it in books in the office is because you naturally become very biased and you're not an objective tester. Right. The answer to that always has to be some combination of anything you can do in terms of an automated system, scripting and dynamic testing. And then to your point, we do have external testers companies. There are companies that you can hire that sort of call upon people to test this and generate reports of things that work and things that don't work. So to your point, you have to do that and you have to do it with people that have not played around with it 100 times.
35:05Otherwise, they tend to be too biased and not really reflect what a typical customer would interact with this like. On the automated side, that is another very active space of applied research and engineering. Obviously, the general approach is to take, again, high-performing LLMs that have been engineered right to play essentially both parts, the virtual customer and the virtual agent, and try to capture as much as you can through that interaction. Of course, voice, again, injects another layer of complexity because text only is very polite and clean turn taking and so forth. And voice, you have interrupts, you have everything else we talked about.
35:50So we have some ways of automating that. But again, always rely on a complementing actual sort of mechanical Turk-like, if you will, testing that's independent. Otherwise, you're right. You're bound to discover the bugs when it hits the customer, which you don't want. And so on the external testing firms, they have armies of freelancers, I imagine, out there, sort of like Mechanical Turk, as you said, who, you know, ask, interact with the LLM. And then all of those interactions are captured and collated. and then you can see where you're weak, where you're strong, and that sort of thing. Do you have an LLM that's playing the customer?
36:43Yeah, you have both. And you ask them to interrupt each other at times with some probability, and you ask them to maybe ask questions off topic. And, yeah, there's just the whole rich world of how do you design these simulated customer and virtual agent interactions, which, again, helps you oftentimes discover many of the bugs or the issues, but it's never 100%, so it's always highly recommended to complement that with you. And are there companies that provide that service? They have conversational LLMs trained for evaluation of other LLMs And I would imagine you set it up and let it run for two days or however long, and it just accumulates this body of data that then you can analyze.
37:43Is that third party or done in-house? So it's done in-house. I wish there were a third party. That might be a good idea for a business down the road. It requires linguists and ML engineers to put that together. And at this point, we did all that in-house. And how long would you let a system run against itself like that or in that evaluation? Yeah, usually it would entail tens of thousands of virtual conversations to get a good signal as to whether. Yeah. So what are the issues that you're able to capture that way? And then real conversations at that point could be in the order of hundreds, which doesn't take that long to run through.
38:25As you mentioned, you have recordings, you have transcripts, and you can learn what worked and what didn't that way. Yeah. And is the system, once it's live, is it training on interactions as it goes forward so that it's continually refining itself, discovering new corner cases? Yes, but not in real time. So you do run it and you collect information and you periodically use that to improve the system after running exhaustive evaluations and testing to make sure that you've improved what you thought you've improved, but you didn't break anything that was working before. And then, yeah, and then you sort of have the next release.
39:10So it's kind of there's a cadence of these updates that does build on, of course, real conversations that are being collected. The other thing maybe worth mentioning is that obviously there's a cost reduction value proposition to all that. But beyond that, and I remember that was strongly the case with McDonald's, what is really exciting about this kind of solution is that it can really improve the customer experience. I mean, not only are you not going to wait 23 minutes on the line sometimes until a person picks up, which is not great for the customer, not great for the brand perception, but you can do things like A-B testing.
39:42And companies sometimes want to say, well, we want to respond this way to a question, not that way. And it's a challenge to convince 5 ,000 people somewhere in the world to do so tomorrow. It's much easier to do that with machines. Analytics, these are the kind of things that with humans, it's just objectively very challenging to do. And these solutions offer so much to our clients, the enterprise, but also to the customer, him or herself. So it's really transformative. Yeah. So you guys started in 2022.
40:19You're not very far into this, but we're not very far into the generative AI revolution, if you want to call it. And looking back, I mean, to what you were doing before Gen.AI, you know, a lot of those solutions, I imagine, look kind of crude or primitive from today's point of view. What's your expectation in how your business and product will evolve in the next five years? and what do you think will change? Yeah, so we're very bullish and optimistic about the field. I suspect, you know, it's not quite clear whether it'll continue to advance in the pace that it has last couple of years, but there's so many people working on this stuff that it's hard to believe it'll slow down anytime soon.
41:19I think all the pieces will get better. I'm not sure how familiar with the latest text-to-speech, the speech synthesizers. There's many companies, including, of course, some of the big ones that have offerings. And it's gotten to the point where it's really indistinguishable. You hear somebody you know whose voice has been cloned saying things that they've never said that the untrained ear certainly can't pick up on it. So I think things like that will continue to improve and become really indistinguishable. The speech to text has continuously improved, and that's going to continue. From our end, there's interesting challenges.
42:02Frankly, some of what it takes to get voice AI to work well is almost AGI-complete problems. But we think that there'll be improvements. One of the things we're working on is trying to train LLMs to not just predict the next word or the next token, which is generally how it's been done so far, but try to look over some horizon and optimize over that. So slightly less myopic optimization, if you will. We're not the only ones working on that. Humans do that very well. And then some of the criticism that's been leveled at LLMs is that they do a lot with this trick of predict the next token. But what if you're able to, like we do, kind of look a little bit further out and produce a token or an action that maximizes some return over some horizon?
42:50So obviously reinforcement learning is all about that. It's about solving what's known as the credit assignment problem. If you have some horizon, something good or bad happened, can you go back and attribute credit to it so as to try to repeat things that work well and avoid things that didn't? So I think all that's going to enrich LLMs. And it's not just going to be language. As you probably know, there's a lot of these multimodal models coming out now, and they're all going to be stronger and better. I think some of the challenges have to do with implementation, the GPU shortage. There's interesting papers very recently that asked the question, you know, can some of this be implemented in CPU, right, which will cost a lot less and reduce the shortage problem around GPUs.
43:38So I think those are some of the exciting things that are over the horizon. Yeah, yeah. Yeah, yeah. It's really a remarkable time. Hi. Good tech solves problems that you've thought about. great tech solves problems that you haven't even thought of. What can the commerce platform trusted by millions of merchants do for you? It's time for Shopify, the commerce platform revolutionizing millions of businesses worldwide. Whether you're a garage entrepreneur or IPO ready, Shopify is the only tool you need to start, run, and grow your business without the struggle. Shopify puts you in control of every sales channel.
44:20So whether you're selling satin sheets from Shopify's in-person point-of-sale system or offering organic olive oil on Shopify's all-in-one e-commerce platform, you're covered. Shopify powers 10 % of all e-commerce in the United States, and Shopify's truly a global force, powering Allbirds, Rothy's, and Brooklyn, and millions of other entrepreneurs of every size across over 170 countries. Plus, Shopify's award-winning help is there to support your success every step of the way. Sign up for a$1 per month trial period at shopify.com slash IonAI. Give them a try. They support us, so let's support them.
45:09That's it for this episode. I want to thank Itmar for his time. If you want to read a transcript of today's conversation, you can find one on our website, IonAI. That's E-Y-E hyphen O-N dot A-I. In the meantime, remember, the singularity may not be near, but AI is changing our world. So pay attention.
From the publisher
Join host Craig Smith on episode #166 of Eye on AI as we sit down with Itamar Arel, CEO of Tenyx, a company that uses proprietary neuroscience-inspired AI technology to build the next generation of voice-based conversational agents.
Itamar shares his journey that started with academic research to becoming a leading tech entrepreneur with Tenyx. We explore the evolution of voice AI in customer service and the unique challenges and advancements in understanding and responding to human speech.
We dig deeper into Tenyx's unique approach to AI-driven customer service while exploring the production and design considerations when developing cutting-edge AI technology.
Make sure you watch till the end as Itamar shares his vision on how AI is going to reshape industries and help advance modern day businesses.
Enjoyed this conversation? Make sure you like, comment and share for more fascinating discussions in the world of AI.
This episode is sponsored by Shopify. Shopify is a commerce platform that allows anyone to set up an online store and sell their products. Whether you're selling online, on social media, or in person, Shopify has you covered on every base. With Shopify you can sell physical and digital products. You can sell services, memberships, ticketed events, rentals and even classes and lessons.
Sign up for a $1 per month trial period at http://shopify.com/eyeonai
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview and Introduction
(03:35) Itamar Arel's Career and Introduction to Voice AI Development
(07:39) Key Differences in Current and Past Technology Solutions
(09:10) Advancements in Voice AI and Large Language Models
(11:00) The Inception and Evolution of Tenyx
(13:29) Challenges in Developing Voice AI
(18:27) Innovative Approaches in Voice AI Development
(21:05) Data Handling and Fine-Tuning in Model Development
(25;41) How To Standout In The Crowded AI Market
(29:44) The Future of Voice AI and Generative Models
(37:14) Testing and Evaluation of Voice AI Systems
(40:10) Where Will AI Be in 5 Years?
(43:42) Closing Remarks and A Word From Our Sponsors




