In short
Podcast Summary: The Real Work of AI Translation: Frameworks, RAG, and ROI
Podcast Overview Title: Talking AI Host: Matt Paige Guest: Olga Beregovaya, VP of AI at Smartling Description: This episode explores the complexities of deploying AI-powered translation systems at scale, focusing on Smartling's innovative solutions and the broader challenges within the field of AI translation.
---
Key Topics Discussed
- Daily Life of a VP of AI (02:39)
- Olga shares insights into her daily responsibilities, emphasizing the need to stay updated on AI developments.
- The rapid advancements in AI, particularly in model releases, have significantly transformed the landscape.
- Navigating the AI Landscape (04:04)
- Importance of filtering through the noise in the AI discourse.
- Need for practical experimentation with AI tools rather than solely relying on external information.
- Smartling’s Model Evaluation Framework (04:55)
- Smartling employs a model evaluation framework to assess new language models.
- Framework includes automated metrics (F-measure, BLEU, METEOR, COMET) for swift evaluation.
- Evolution of AI in Language Processing (08:50)
- Transition from rule-based systems to transformer models.
- Existing challenges with language translation and the need for human-in-the-loop solutions.
- Challenges in Language Translation (15:38)
- Complexity of translating diverse languages.
- Examination of linguistic nuances and the implications for model performance.
- Enterprise Scale AI Solutions (21:38)
- Discusses the myth that simply integrating AI models like GPT can resolve enterprise translation issues.
- Importance of adaptability and multi-model approaches in meeting varied enterprise needs.
- Future of Language and AI (28:52)
- Speculation on a universal language and the continuing significance of diverse languages.
- Prediction that language families will persist, while written and spoken modalities may evolve.
---
Key Takeaways
- Model Evaluation Framework: Smartling's approach emphasizes systematic testing of language models to ensure quality and efficiency.
- Complexity of Language Translation: Different languages present varying levels of complexity for AI models, influenced by training data and linguistic structure.
- Scalability of AI Solutions: Effective deployment of AI in enterprises requires careful consideration of language coverage, input parsing, and latency management.
- Future of Language: While technological advancements may bridge communication gaps, the richness of language diversity is likely to continue.
---
Notable Quotes
- "The daily life has changed dramatically... staying appraised is one of the most important things."
- "Don't fix what's not broken. Don't pull in Gen. AI where good old rule-based deterministic approaches work just as fine."
---
Additional Resources
- Smartling Website: [Smartling](https://www.smartling.com/)
- Olga Beregovaya on LinkedIn: [Connect with Olga](https://www.linkedin.com/in/olga-beregovaya-04b5/)
Mentioned Tool
- AI Opportunity Finder: A free tool by HatchWorks to uncover high-impact AI use cases tailored to businesses.
---
Conclusion This episode of *Talking AI* provides valuable insights into the challenges and strategies involved in deploying AI translation systems, showcasing the expertise of Olga Beregovaya and addressing the broader implications of AI in language processing.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00First of all, there is this urban myth, right? I'm just going to plug in GPT and solve all of my enterprise translation problems. And that's where you go, yeah, but the models don't quite self-police and self-heal yet. So how are we going to catch your model hallucinations and mitigate your model hallucinations? Welcome to the Talking AI Podcast, where we talk AI with both experts in the field and early adopters. I'm your host, Matt Page, and we're here to demystify AI for you so you can get some value from it. Let's talk some AI.
0:34Language is the original interface. Now with AI, we're rewriting how humans communicate across borders, platforms, and cultures faster than ever. But while AI translation is getting better by the day, actually deploying it at scale across dozens of languages, format systems is a whole different challenge. And this applies to deploying any AI solution at scale. Today, I'm joined by Olga Bergovia, VP of AI at SmartLink to break down what it really takes to reach human level translation and where LLMs shine and where they fall short. Welcome to Talking AI, Olga. Thanks so much, Seth. Thanks for having me.
1:09Excited to get into this chat. And I'm curious, let's just set the stage here. Give me some context, A, for what SmartLink does. But I'm also curious with my guests that are in these roles of VP of AI, what does your day in the life look like as VP of AI nowadays? Give us kind of that breakdown of your role. Okay, so first of all, Smartling. Smartling is a fully AI-powered global content delivery platform, or to say it's simpler, translation delivery platform, which equally can deliver fully automated translation and will touch on human parity later, or equally provide facilities for human-in-the-loop where human-in-the-loop is needed, both on the prompting side or design side and post-editing side.
1:51So that would be Smartling in a nutshell. I guess that's like global businesses that have exposure all over the globe and you have this need for multi-language, whether it's marketing or any number of things. I'm assuming that's kind of where y 'all come into play. I would say we come into play if somebody looks at our website. We have a pretty substantial, like a whole page, huge page of tiny logos, list of clients. And they really do come from all ways of life. And you're absolutely right. servicing a number of domains. Obviously, legal will have a certain set of requirements, right? Retail will have a completely different set of requirements.
2:29Online opinion portal is a piece of its own. So what does it consist of? It consists of linguistic particularities of each content type and it consists of the input format. And again, to your point, what you would pull in from them would be dramatically different from what you would pull in from Git, right? In terms of the string formalism and what the content would look like. So there are multiple dimensions to it. And then we add 7 ,000 languages spoken in the world and that would pretty much give you a picture of what we do. I wouldn't say we do all 7 languages though. Yeah. But just in general, that's amazing that there are 7 ,000 languages spoken in the world.
3:08That still kind of blows my mind. There's that many different languages out there in the world. So I'm curious, your role as VP of AI. What does that look like on a daily basis? And I'm curious, how has that changed since the advent of generative AI, chat GPT? Because AI has always been cool and popular. It's at a whole nother level right now. What does that look like for your daily life and work? The daily life has changed dramatically. Basically, one needs to be up one hour earlier just to go through all the TLDRs and all the latest model releases. And obviously, in the modern day of models battle, right?
3:46And each model tends to also evolve and come with its own flavors. So it's just staying appraised. Staying appraised is one of the most important things. And it has never been this dramatic and this massive ever in my 25-year-old career in language AI. So first of all, it's just being in the know, looking at articles published, and just you wake up to a new day every day, both in terms of model number, anything, number of parameters, capabilities, whatever, everything. So it's a new world. Second is not to fall for this trap and actually navigate your way and research and development way and company's way through what matters and what can be parked or would be irrelevant to our work.
4:32So get as much knowledge as you can and steer course as much as you can. Yeah. And that's a great lesson, I think, for everybody listening to like the first part you were just hitting on. It's like, how do you get identify the signal through the noise? And it's so difficult. I had a strategy at Hatchworks AI and I feel the same pain because every day it's like some massive announcement that typically you would get a few a year. It feels like they're daily or weekly right now. And I think the point you mentioned is like, how do you get through the hype? And it's difficult because everybody's riding this wave right now.
5:08They're intentionally playing into the hype. And for me, it's I just as much as I can. And I think for people listening to go try the tools yourself. I think that's really important because if you're just learning by listening to others, that can be difficult to figure out that signal through the noise sometimes. Quick break in the pod. If you're listening to this podcast, chances are you've been thinking about how to actually use AI inside your business. And that's exactly why we built the AI Opportunity Finder. It's a free tool that helps you uncover high impact, tailored AI use cases based on your business, your goals, your pain points, and your industry.
5:45No fluff, no generic use cases, just real ideas that fit your business and the ranked by ROI potential. It takes about three minutes to run and it's like having your own personal AI strategist for free. If you want to try it for free, check out the link in the show notes or go to hatchworks.com backslash AI dash opportunity dash finder. Oh, absolutely. And to that effect, filtering through the noise, I think we at Smartling made a solid investment decision to having a model evaluation framework. So take me through that. That's interesting. If we were to test every little model that's out there and we have this research paper and news Slack channel on my team.
6:28And of course, we drop every time Hugging Face announces something or there is a new announcement. Of course, we drop information about this model. but it's more of an FYI. We usually identify the models that are fit for our purposes. We usually already have a pretty good idea just by looking at model parameters, just looking at model card. We would know whether it's worthwhile for us to look at it. But then we have the evaluation framework where we have the usual suspects that like, I don't know, F-measure or blue, meteor, comet, and like basically a set of automated metrics that we would just run the model through and we would get a pretty quick answer on shall we proceed or are we good where we are?
7:13And again, so someone would - Yeah, go ahead. Is this almost, so it's grading essentially models, new models that are being introduced to determine, okay, do we shift the model here or there? And you've almost built this automation into your system and your process. Is that kind of how to think about it? That's kind of how it is. It's, yes, we have the automated system. We produce translation and what is there to translation? There's quality of translation and there is ability to estimate quality of translation. And then feedback loop from there and deliver better and better translation. So we're not going to, we know what we do as a platform.
7:48So for us to build this evaluation framework is fairly easy because we know that the model needs to be fit for purpose. Right. The model needs to deliver to our objectives. so that's why the evaluation like we don't need to evaluate how the model calculates right we don't need to evaluate how the model i don't know dream something up summarizes e-discovery email we have our set yeah because you would go on forever it would be a million since our test cases are very we know exactly what the test cases are both on the coding side of the house and on the linguistic delivery side of the house we run it against our test cases it's very easy.
8:27It's saved out, worthwhile. And we're going to get to this later, but that's a key thing, I think, for enterprises that are looking to leverage AI embedded into their systems. Like what Olga just went through is a great approach to staying up to date because a lot of people build something, they just assume, oh, I built it. It's good to go. We got ChatGPT, whatever model in there. Things change. Being adaptive is very important to stay up to date in a sense. And And maybe there's another dimension to it. By now, I don't know, something very dramatic needs to happen, like a major breakthrough for us to pivot from our models that we've selected for our processes.
9:06So if we know that Claude is like Sonnet is going to work somewhere, but 3.5 is a legacy model that tells us, okay, hey, these enhancements happen at four. But by then, we already have our Sonnet prompts. We have Sonnet integrated. I mean, we have it already integrated with our systems. So one other dimension is something super dramatic needs to happen. There needs to be like crazy good at translation, quality estimation, summarization model. Then we're like, okay, we'll pivot. Other than that, we have our model garden, whatever the best word is, stealing from Google. And we try to still confine to Bedrock, Azure, or OpenAI.
9:44So basically, there is a set of models that we choose to work with. and then we would wait within generations of those models and different sizes of those models. Exactly. That makes sense. Like the juice has to be worth the squeeze to go through the effort of changing something to get the incremental benefit. Otherwise you're just constantly spinning in circles, I can imagine. And what's I think interesting with SmartLinks, the company has been around for 16 years and I was looking over at your LinkedIn, you joined in September of 2022, which is right before like ChatGPT came out, this like seminal moment, even though Transformer architecture has been around for since what, 2017?
10:23It just kind of took that moment. What does this evolution look like where we're going from very deterministic systems to now we have these, this new tool, this probabilistic system? What does that look like Like just in your career as being in the AI space and now you have this new like amazing thing to leverage. So I joined SmartLang two and a half years ago. But by then we already had a chance to play with transformer models in neural machine translation. And GT2 was already customizable. So it already kind of gave us a pretty healthy preview of where things were going. So that was the right time to join the company that specializes in AI driven translation technology.
11:03but if you look back at my career you're absolutely right okay first things first large language models have a language component to them so to an extent as much as foundational generalized models can do everything under the sun right language has always been a piece of that so that's statement one statement two you don't necessarily go from deterministic approaches to transform models, you can also augment. And that's one of the principles we follow at Smartlane. Don't fix what's not broken. Don't pull in Gen.AI or any other form of machine learning where good old rule-based deterministic approaches that are significantly cheaper work just as fine.
11:51Like things like, for instance, in French, you want to have a space before the slash and after the slash. You don't need an LLM to do it for you. Right. You write a simple rule and just say, so I think it's augmented approach. It's not necessarily transforming, but go ahead. That was just a great point there too. Cause like, I feel like a lot of people fall into this shiny object syndrome where they're like, oh, I have this new shiny toy. I got to go apply it to everything, but it doesn't necessarily need to be applied to everything. So I think that's such an important spot when you're thinking about how you adopt AI.
12:27it's okay where can it best be applied to current things that's one and i think two it's like where can it be applied to things that were never possible before that now this enables and the stuff that's fine working as it is like to your point just leave it alone it's fine it doesn't need to be applied to everything yeah but to answer maybe a previous question about the trajectory in nlp space the evolution was very just like probably everywhere else of machine learning, right? You go from rule-based, right? You go from rule-based and machine translation would be a great example from that. And for a while, for a long while, rule-based automated translation was dominating the scene.
13:06So that's where I come from. Again, having been around the industry for 25 years, even before rule-based machine translation, you had approaches like using Levenstein distance to identify fuzzy matches in previous translation. So it was always, I would say that language industry, and that's the beauty of it. By now, we might be even more ready for the AI revolution than maybe other segments and other sectors are. Because we've always worked with some language-related processing tools. And now in LLMs that are, yes, there is multimodality, but mostly LLMs started as text-based models. I think we were going through just pure math to rule-based, then to statistical approaches.
13:50and now finding ourselves in transformer models, right, and all of their iteration, that was a very natural trajectory following, this is what's new in ML and this is how it applies to language. And I've been watching it throughout my career. Backend changed a lot, but even more, the user experience changed dramatically. And what translators can and cannot do, you actually need to rewrite and retrain your translators and your linguists every time when you approach surfaces. example statistical machine translation does not hallucinate nmt took to hallucination immediately we all know about large language model hallucinations so suddenly a human in the loop is dealing with completely different bag of cats whatever completely different set of issues cana worms what do you have in fact i like bag of cats though i think there's cat Hold on, sorry.
14:45Here I am preaching that English is not my first language, so that can be forgiven. Yeah, again, I've been watching it and being a part of it throughout my career. But the explosion that we're dealing with now and the realm of possibilities in natural language processing that we're exposed to now is nothing like what we've seen before. What is your first language? Russian. Russian? Yes, Russian. How many languages do you speak out of curiosity? Because I come from St. Petersburg, we were very close to Nordic countries. So I happen to speak, I'm native fluent in Danish, I would say, reasonably decent when I have practice in Swedish and Norwegian.
15:22And because I spend half of my time in Mexico, there is like some twisted version of Spanish that I speak. But as you read academic papers on structural linguistics, you natively have to learn how to read French and German papers. Because a lot of traditional school of linguistics comes from Germany and from France. Yeah, but you have a nuanced understanding of language in general, which is pretty cool being in this space. My language expertise goes into the very minute bit of Spanish, and that's about it. No, it's also because I studied structural linguistics. I mean, all the syntax morphology, phonology, and language structure.
16:00Really? Oh, cool. field of study and everybody on my team, which is essential for being in that field. Everybody on my team has understanding of structure of languages. We do not, we would go again, 7 ,000 word languages, 7, 71, right? Or whatever that is. Say out of them, we can service 500. I cannot have a team that speaks Swahili, Punjabi, Yiddish, and pick your fourth language, whatever you want, Estonian. Yeah. But it's essential to understand, okay, here you are in the Finno-Yukuric family, or in family of romance languages, it is essential for the team to understand how languages operate and what you can get out of Chinese and idiographic languages, right?
16:42As opposed to agglutinative languages. So that's the essential knowledge for our field altogether. And I'm blessed to have a team that understands it and can work with it. That nuance is really interesting. I'm curious on that same thread, what makes certain languages easier or harder for an LLM to understand, predict, interpret, whatever term you want to use. Do you notice any nuance there between maybe these different language families and whatnot? Or is some harder to translate between others, maybe? I mean, first of all, if you look at, and I wouldn't go, I'll probably go a little bit, roll back a little bit and go to transformer model because neural machine translation equally is a transformer model, just with fewer parameters, right?
17:26But, and purpose built, But other than that, we're still talking about transformers. First of all, just because of the nature of the training data, and if you say the model has scraped all of the Internet, we know how much English is in the training data. So first things first. That's a great point. And I mean, I can go for another hour about training data not being diverse enough. And sometimes you would have the lexical coverage, but you would not have the cultural phenomenon. And that's when you start producing. The nuance. And then you start producing absolutely weird things in the target language that look like the target language, smell like the target language, just have nothing to do with the target language.
18:07So that's one. But because the majority of training data, unless you fine tune for a specific language, the majority is English. Having English as a source is the ultimate dream. Right. Right, because your input is extremely easy to understand, right? Your input is super understandable and the model would have no issue understanding the input sentence. Just because the English is the prevailing language of those models. Now, we actually run a grid of language complexity or automated translation. And there are some language families that render themselves equally good to translation with large language models and with neural machine translation.
18:51and that's Romance languages and Nordic languages. If we want to show off to a customer, if you really want to put your best foot forward, grab Brazilian, Portuguese and shine, right? Vocabulary, grammar, reasonably, fairly simple. Training data set, huge. Same for Latin American Spanish. Yeah, training data set, grammar of the language. So all of that, between all of that, I would say that Romance languages, also the way you tokenize Romance languages probably make them a very easy translation target. But then in the bottom of that grid, we would have Finno-Ugric languages like Estonian, which would be extremely complex grammar, right?
19:33Of Finno-Ugric languages, Finnish or Hungarian, combined with lower training data set, lower language coverage. And this is where you would have complexities. Now there are some interesting edge cases like Southeast Asian languages, which actually render themselves, If you have enough of training corpora, because of the nature of the languages, it's actually, they're fairly easy to translate, even despite the give or take sparse training data set. It's interesting, too. I saw this study, and I think it was related to Chinese, Mandarin, maybe. Because the language itself is character-based, and those characters have symbols, versus English and every other language, it's based on an alphabet.
20:13I forget if it was cheaper, easier. be a combination of both for training on that. I guess having the corpus of training data is a whole other element, but it's an interesting kind of experiment they were testing there. And LLMs think in tokens, right? They understand tokens, right? And if the token boundary is very obvious in a specific language, then obviously it renders itself much better to translation with LLM. That's a good point. I mean, it's worth a dissertation on its own. We were hitting a wall, but I think we nailed it. there's a funny thing if you think about UI user interface quite often you have incomplete sentences or they're not sentences of any kind at all but yet whatever the translation is you need to fit it into a specific box right because you only have this thing and this is where you go I have a word here a word sentence here and yet I have my counterpart the system that only understands tokens and this is where okay I'm going to give you a chopped off string that makes no sense whatsoever and that it doesn't even doesn't even contain real words and this is like the balance between what's a word and what's a token is at play for many languages much more than it is for other yeah i remember dealing with that way back in the day before all of this crazy llm stuff and we had humans doing all this but i remember that same consideration around the ui because some languages are just naturally longer and changes the the whole design and feel and everything there when you do this translation.
21:45That's bringing back some memories for me. Yeah, but there are two ways that we can expand the box or shrink the text. They're only one of the two. This is one of the examples where LLMs do not shine by any means. So there are a lot of bells and whistles that you need to add. So there are cases like that. In that same two, and this was maybe a year or even maybe two ago, earlier days, I remember them doing some type of test on different languages with the cost of AI, like a per token cost, and it being more expensive for certain languages, whether it was based on training data or length of the language and the output.
22:21So it was almost like this introduced bias, in a sense, based on the language you were interacting with, which I thought was kind of interesting. Yeah, it is very interesting. But again, it's what is a word, what's a token? Because if you burn all this money and all this effort, like whether you host a model or not, if the token is equivalent of a word and it's shorter, right, The word is shorter than inherent can be cheaper as opposed to longer. If you look at German compounds for example. Yeah. So yes, the whole fine tuning and inference cost is, I mean, again, it's a thing of its own. Let's get into that.
22:51Let's shift to this whole discussion around enterprise scale, because anybody nowadays can go build something in lovable, connected to AI or a custom GPT and all this type of stuff. But actually building a system that has enterprise scale has a whole nother list of things you need to consider. What things do you think about when you're going to enterprise scale in a sense? Having said this actually enterprise ready, what are those vectors you think about those attributes when you're going through that process? Well, among other things, we think about vectors. But I mean, that aside. Yeah, vectors realistically too.
23:29I feel just too good to miss.
23:34so first of all there is this urban myth right i'm just going to plug in gpt and solve all of my enterprise translation problems and that's where you go yeah but the models don't quite self-police and self-heal yet so how are you going to catch your model hallucinations and mitigate your model hallucinations so you can just think about the quality of translation. A, what are the language-independent approaches that would allow us to service different content types across a variety of languages? And you identify things like RAG, for instance, right? Enterprise translation, enterprise-scale translation systems.
24:15We store translation memories, right? We store or extract style guides or stylistic preferences. We store glossaries. We parse multiple file types. So first things first, We want to make sure that our customers get the advantage benefit of being able to parse and prepare for automated translation different inputs. So that's one. Second, and that's something I've noticed working with a lot of buyer side enterprises, often engineering teams think about English. A lot of testing is done in English. Sometimes there would be like Japanese or an Indian coder on the team that would say, okay, look, I just quickly tested it for whatever.
24:55for, again, pick a language for Japanese. And it's very important for scalable enterprise implementation, global enterprise implementation, to make sure that you have a solution that works across a variety of languages, that you're not trapped in monolingual or bilingual universe and you address the needs of all the geos. So language coverage, breadth of language coverage, ability to ingest from different content types and parse them out. Models don't play well with special characters to date. Models don't play well with emojis. Models freak out when they see nested HTML tags. So obviously, when you take this into consideration and they change again for different languages, you need to be sure that your AI approach is fully internationalized from the get-go.
25:48So then, there we are. They're your hallucinations and you need to have hallucination mitigation mechanisms that are language neutral, universal. So you can phase in languages and phase out. Multi-model approach is also important because of language coverage. Because different models process different languages differently. So like Smartling, for instance, we, as I said in the beginning, we're not stuck with a single model, a single family of models. But we know Llama would do great here. Gemini would do great here. OpenAI would have a good coverage, but come with its own set of challenges. And once, say, you've solved it all, you have your reg mechanism, there come your inference time and latency.
26:33Right? Yeah, which is important. And if you just have one endpoint and that's it, before you know it, you are stuck. And if you need to publish something real time, you really need to have the latency mitigation mechanism that can follow different avenues, right? It can be multi-threading. Yeah. And for the audience, when you're talking about inference time, computing, things like that, if you're not familiar with that term, those listening, think of when you're interacting with ChatGPT, how long it's taking to give you a response back that's dealing with that latency piece, which can be very bad on the user experience side if it's taking forever to respond based on the use case you're using it for, essentially, right?
Read the full transcript
27:11Yes, absolutely. And actually one fun fact, I was just talking to the head of R &D in my department and we were talking about why we're not playing enough or as much with latest generation of reasoning models. And the answer is very simple. If something needs five steps to think, I don't have the time for those five steps to take place. So they maybe could even produce better results on quality assessment. But until we figure out how to deal with the delays, we probably would try and stay with LLMs versus going to. And that's all about tailoring the model to the use case, too. That's a great example of the latency that attribute or vector there, in essence, is prohibitive to using it for that particular use case.
27:59And like when you're considering models, like obviously there's cost performance. Are there other attributes you consider when you're thinking about a model or is it primarily those? Our engineering department runs reports capturing signals on latency. So that obviously that's one. Latency. Yeah, it's another one. Latency. So that's a dimension. Do we, how often do we need to default to the fallback? And that's actually another part of enterprise strategy. Like how often would a model choke on something or fail for one reason or another? How often do we need to invoke the fallback model? So that would be a second.
28:34How often does a model say on 10 ,000 strings? How often, how many times did the model fail, choke, interrupt the workflow? So stability, how many times do we need to invoke a feedback? Like the model is just not performing in engineering terms, latency time, language, quality of translation. And then quality of quality estimation, how good is a model at predicting the quality? So language coverage is another parameter that we would definitely take into consideration. Model size. Model size. Can you get the same result of medium-sized model that you can from, for instance, you look at Amazon Nova family, for instance.
29:14Where can we get with Nova Mini and can we get to the same place as Nova Pro? obviously reporting to the ceo with a cfo in the cfo in the running uh running the pnl obviously the model cost and the different yeah model cost the roi is definitely one of the huge factors but that also constraint one of the smart link principles is constrained race creativity So that makes us prune our prompts. Like if you look, in order to get great translation, you've just burned through 3000 tokens and the CFO is not going to be happy. That's where we need to sit down and figure out what we can cut out, what we can rag in and what's not necessary at all.
30:01Exactly. Exactly. So I'm curious, like completely pivot, like forward looking. Do you think there's a point in time where what's happening technologically today and it's almost like this accelerated globalization in some ways, do you ever think we reach like a universal language or do you think having, like you mentioned in the beginning, 7 ,000 languages is something that will just always be true? Do you think there's almost this language extinction that may come to pass over time? Getting, well, that's a bit philosophical. Use your reasoning model part of your brain right now. It's pretty esoteric right there.
30:40Let me think about it. The way I like to think about it is more erasing the linguistic divide, right? That's in place right now and enabling communication between different languages, language family. Like in the past, you did to pivot through English. I have a pivot language for everything, right? Now you can go between whatever the language combinations. so I mean Esperanto failed miserably right so we're talking about that singularity powered by the next generation of Esperanto I wouldn't think so I would think that what's going to happen is probably there's so many parameters to it I mean will long-tail languages go extinct we've been speaking about it for a long time what I think is going to happen is the existing language families people still talk people still communicate you still have the spoken modality i think what's going to happen though the spoken modality might get simpler but will still retain its current qualities and the language families are going to stay exactly what they are now when you look at the written modality spoken you and i talking right yeah the way you and i talk probably is going to be more preserved than the written modality that's my prediction that's interesting There's something you said that just triggered an idea in my head.
31:59Like you mentioned the connector piece. What instantly came to my mind is like APIs in a sense to where it almost enables more communication, but in your native language. So me speaking to somebody else in Mandarin, that becomes easier with this technology. It doesn't necessarily mean I'm changing the language I'm speaking in, but it allows more nodes of communication in an easier way. I think that's a really interesting notion, which doesn't lead to extinction. Yeah, nothing goes extinct. What's going to happen, though, and there is a lot of research about it, we always say, oh, we need human corpora to train models.
32:43The reality is, over the past five years, the way we write and we speak is so impacted by the way the models, like machine translation or generated text, we are influenced by the models just as much. So it's kind of funny, it started going both ways. example yeah you'll learn if you're not a prompt engineer yet you want to have a nighttime philosophical debate with your chat gpt you kind of know how to phrase and how to formulate your ask to get the most relevant results so thousand asked in the way you communicate has changed dramatically so it's not that there'll be universal language but more i think there'll be this universal communication mode, universal communication modality, where we'll know how to interact with models.
33:36Yeah, that's interesting. I think that's a great stopping point right there, Olga. Thanks for being on. Where can people find you, find SmartLink, learn more about it? Hit us with the websites and everything where they can find you. Finding me is simple, adding me on LinkedIn. That's pretty much where, yeah, finding me on LinkedIn. And in terms of Smartling, first of all, we're very diligent about publishing blogs and publishing our materials on our website. So that would be Smartling.com. And again, we're super diligent in our social media presence. So following us on social media platforms, we're very active on LinkedIn.
34:11We publish something pretty much every other day. So LinkedIn and Smartling website would be the best ways of equally finding me and finding the company. Awesome. Thanks for joining me and talking a little bit of AI. Thanks for having me. Thanks for listening to the Talking AI Podcast. If you enjoyed the show, give us a follow or subscribe on your favorite podcast platform. And don't forget to leave us a review. We love those. For more info on Talking AI, visit TalkingAIPodcast.com. The single biggest mistake we see companies make with AI is they don't properly train their teams. We see it all the time.
34:51Companies roll out AI tools and expect people to just figure it out. But using AI effectively requires a totally different mindset and skillset. And that's exactly why we built training for every level of your org, from AI training for teams and executives to training engineering teams on our generative-driven development methodology. Or if you've already identified your AI use cases and want to just prioritize where to start, we offer an AI roadmap and ROI workshop to help you build a quick plan. It's all about going from we should use AI to actually driving real value with it. Head over to hatchworks.com to learn more.
From the publisher
In this episode of Talking AI, Matt Paige speaks with Olga Beregovaya, VP of AI at Smartling, about the complexities and challenges of deploying AI-powered translation systems at scale.
Olga provides an overview of Smartling's platform, which leverages fully automated and human-in-the-loop translation solutions, and discusses the evaluation framework they use for selecting suitable language models.
The conversation delves into the evolution of language processing from rule-based systems to transformer models, the impact of generative AI like ChatGPT on daily workflows, and the intricacies of translating diverse languages.
They also explore the potential for a universal language and the importance of adaptability in AI deployment within enterprises.
--
Key Moments:
- 02:39 Daily Life of a VP of AI
- 04:04 Navigating the AI Landscape
- 04:55 Model Evaluation Framework at Smartling
- 08:50 Evolution of AI in Language Processing
- 15:38 Challenges in Language Translation
- 21:38 Enterprise Scale AI Solutions
- 28:52 Future of Language and AI
--
Key Links:
Mentioned in this episode:
AI Opportunity Finder
Feeling overwhelmed by all the AI noise out there? The AI Opportunity Finder from HatchWorks cuts through the hype and gives you a clear starting point. In less than 5 minutes, you’ll get tailored, high-impact AI use cases specific to your business—scored by ROI so you know exactly where to start. Whether you're looking to cut costs, automate tasks, or grow faster, this free tool gives you a personalized roadmap built for action. 👉 Try it now at https://hatchworks.com/ai-opportunity-finder/
