In short
Episode topic: State of the art in AI translation and evaluation, centered on the latest WMT (Workshop on Machine Translation) results, especially the unified metrics/QE shared task using ESA-style human judgments and harder test sets.
Guest backgrounds
Tom Kocmi is a researcher at Cohere on multilingual MT evaluation; previously led evaluation at Microsoft Translator and has long WMT involvement. He helped pioneer the trainable METEOR metric (early 2000s) and later worked in industry, including Amazon and Unbabel (where COMET was developed). Alon Lavie is a long-time CMU professor (Language Technologies Institute) focused on MT and automated evaluation; he has led WMT metrics/QE efforts for years.
Key claims
LLM translation is now fluent and context-aware, shifting the bottleneck toward semantic errors and evaluation. WMT must keep changing tasks because metrics get “gamed” (Goodhart’s law). Many systems underperformed because they were tuned/hill-climbed on COMET, importing its biases.
Notable examples
WMT harder domains included news commentary, literary fiction, speech translation from ASR transcripts, and social/user-generated content. Organizers used LLMs as judges (reasoning LLMs performed best). Human translation did not top the cluster; the hosts caution against “human parity” claims due to tired translators, no pre-translation, and human annotation limitations.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOBackground on WMT Competition
0:45 to 2:00
Overview of the WMT competition and its evolution over the years.
“translation competition, for lack of a better term, and generally the state of the art in AI translation more broadly.”
Guest Backgrounds and Roles
2:00 to 4:20
Tom and Alon share their backgrounds and roles in AI translation.
“Can you just tell us a bit more about your background, how you got into translation slash language AI and your current role, maybe starting with Tom?”
Progress in Machine Translation
4:20 to 6:00
Discussion on advancements in machine translation evaluation and performance.
“Meteor is actually used beyond machine translation, actually.”
Shifts in Translation Technology
6:00 to 10:00
Exploration of the transition from neural to LLMs in translation technology.
“summer and I decided to come back late in my career to academia.”
Understanding WMT's Importance
10:00 to 12:00
Why the WMT competition is crucial for advancing machine translation research.
“But now with LLM and the move to LLM doing the translation tasks, they do it in full context.”
Automatic Evaluation in Translation
12:00 to 14:08
The significance of automated evaluation methods for translation quality.
“And as well, we are also pushing how to evaluate them.”
Machine Translation Evaluation and Quality Assessment
14:08 to 18:06
Learn about the evolving methods for evaluating machine translation quality and the importance of quality assessment metrics.
“You need an automated evaluation system that would actually assign scores and give you that type of information at a granular level as much as possible.”
Challenges in Machine Translation Quality Assessment
18:06 to 22:20
Explore the challenges faced by human evaluators and automated systems in distinguishing translation quality.
“One was the data setup that we adopted from the general shared task, which much larger, longer segments.”
Findings from This Year's Evaluation Task
22:20 to 26:37
Discover the key findings from this year's evaluation task, including the performance of large language models.
“and indeed, it's not the case for every of the language pair that we saw this year, especially with the techniques that Tom mentioned for making the data harder.”
Bias in Evaluation Metrics and Model Performance
26:37 to 28:00
Understand how biases in evaluation metrics can affect the performance of machine translation models.
“And all of those variants of those trained neural metrics did not do very well this year.”
Show all 19 chapters
Evaluating Translation Models and Metrics
28:00 to 29:50
Learn about the challenges and biases in automated evaluation of translation models.
“And you get better Comet scores as you do that.”
Performance of General-Purpose LLMs
29:50 to 32:50
Discover how general-purpose language models are outperforming traditional systems in translation tasks.
“Tom, do you have anything to add to the eval part?”
The Role of Multilinguality in AI Translation
32:50 to 35:20
Understand the limitations and potential of multilingual models in translation.
“You can train in a couple of days on a relatively small cluster and get actually close to the state-of-the-art performance.”
Challenges in Human Translation Evaluation
35:20 to 39:55
Explore the complexities and challenges in achieving reliable human translation standards.
“And then on the evaluation side, again, I think we see sort of the artifacts of that.”
Implementing AI Translation in Workflows
39:55 to 42:00
Learn practical considerations for deploying AI translation and evaluation in workflows.
“When there's so few errors, it's so easy to miss an error here and there.”
Evaluating Automated Translation Quality
42:00 to 44:30
Learn about the challenges and processes in evaluating translation quality systems.
“processes that know how to actually leverage these systems that we basically find are performing well into repeatable processes.”
Insights from WMT on Translation Models
44:30 to 46:40
Discover key insights from the WMT conference regarding translation model evaluation.
“But okay, going back to the original question.”
Frontiers and Challenges in AI Translation
46:40 to 50:00
Explore the current frontiers and challenges in AI-based translation technologies.
“I want to close on a couple of big questions or a series of big questions.”
Future Directions for Translation Systems
50:00 to 53:00
Understand future directions and strategies for enhancing translation systems.
“Like, why are we still in a world in which the LLM that's generating the translation doesn't already know everything that there is to know about the quality and doesn't generate errors to begin with, right?”
Transcript
Automatic transcript. May contain errors.0:00With the rise of the LLMs, we needed new ways how to evaluate them because they kind of catch up with the evaluating metrics and then be new problems.
0:11Florian:Hey everyone, and welcome to a great episode here at SlaterPod. Today on the podcast, we welcome Tom Kocmi and Alon Lavie. So Tom is a researcher at AI Lab Cohere and Alon is a distinguished career professor at the Language Technologies Institute at Carnegie Mellon University, one of the leading universities in language AI has been for decades. So we also have with us Slater's own research analyst, Maria Stassimiotti. So Maria covers language AI here at Slater and I'm very happy to have her on. And we're here today to talk about the findings of the world's most important AI translation competition, for lack of a better term, and generally the state of the art in AI translation more broadly.
0:54Florian:So hi everyone and thanks so much for joining. Hello. Nice to meet you. Cool. So look, I'm trying to give a little bit of background to today's conversation. So it's going to revolve around the WMT competition that concluded recently, and we've been covering it on slater.com quite extensively. So just super briefly, and I will ask you to tell us more about it later on, Tom and Alainz, but super briefly, I think the WMT, it stands for Workshop on Machine Translation and it was initially first held like in the mid 20, like 2006 or something as part of a, it was a chapter at the North American chapter of the Association of Computational Linguistics, but then I looked it up.
1:36Florian:It said in 2016 with the rise of NMT, it kind of became, the WMT became a conference of its own. And yeah, so this concluded recently and you both were instrumental in doing this. And yeah, and generally, so we want to talk about that, but we also want to talk with you about the kind of general state of the art in translation AI. So now let me hand it over to you, Tom and Alon. Can you just tell us a bit more about your background, how you got into translation slash language AI and your current role, maybe starting with Tom? Currently, I'm a researcher at Cohere where I'm part of a multilingual team focusing mostly on machine translation and evaluation.
2:15And before I joined like a year ago and before I was working for four years for Microsoft Translator where I've been leading the evaluation part of the whole pipeline. And so I'm leading the WMT conference or mostly the general empty share task where Alon is taking care of the metrics share task or automatic evaluation. So maybe a little bit about the WMT. it's as you said it's a conference nowadays and it contains multiple share tasks that focus on different parts of the machine translation either the how to improve the performance of the models or how to evaluate them or how to evaluate the metrics and this year we hit the 20 years of wmt conference i've had a pretty long career in in this space and i've been here in pittsburgh all around sort of Carnegie Mellon for just about 30 years.
3:18But I was here a professor already at CMU for about 20 years, the last starting from 1996. My career, my research areas have really focused on machine translation. And over the last 10, almost 20 years, actually, one of my areas has been automated evaluation. of translation quality, primarily machine translation, but increasingly any translation quality that involves automatic translation or human translation or a combination of those. So earlier on, soon after BLUR came out in 2002, I actually initiated a research project that developed one of the earliest metrics that was trainable for machine translation, That was the Meteor metric that was actually used by a lot of people in the field up until for over 10, 15 years, I guess, until the neural era came out.
4:29Meteor is actually used beyond machine translation, actually. I keep being surprised. It's the most cited of all of my research on Google Scholar and on Semantic Scholar. I have immense numbers of citations for Meteor, and it's used in various NLP applications quite successfully. But then in 2009, I had a startup. I spun out from my lab here at CMU, a startup company in machine translation adapted for enterprises, a company by the name of Safaba. and long story short, Safaba was acquired in 2015 by Amazon, at which point I took a hiatus from my academic career. I remained as an adjunct professor and continued to do some advising, but I went for four years into Amazon as a senior manager for machine translation and ran the machine translation research group here in Pittsburgh for Amazon.
5:32And after that, I left Amazon in 2019 and I did two senior AI management, R &D management roles, two years at Unbabel. And then for the last two years from 2023 until earlier this summer, I was the VP of AI research at Frays. And then I concluded my executive kind of management full-time role at Frays early this summer and I decided to come back late in my career to academia. But along the way, at Unbabel, we launched this initiative developing a new neural age machine translation evaluation method called Comet, which was widely adopted and is still widely used until today. I think We'll talk about this maybe a little bit later, but I think in the age of LLMs, new technology is already taking over.
6:34So I don't expect that Comet will be the dominant or the primary metric to be used for much, much longer, but it's done very well and has been very useful for progressing the research for machine translation over the last few years.
6:50Florian:Great. Thanks. That was super helpful. So maybe before we go into the WMT, I just want to take a quick detour into, if you were, or before the kind of LLM revolution were in the field of AI translation, how would you recap the past three, four years since, you know, since these models came online in a big way, maybe starting with you, Tom, and then moving to Allon? I would say, like, maybe over the last five years is the best timeframe. So the main topic that's been changing is like the huge progression in the evaluation, where we change a lot of stuff that we've been doing differently. And with the rise of the LLMs, we needed new ways how to evaluate them, because they kind of catch up with the metrics and then be new problems.
7:35I think another change is mostly about how the translation looks like from the LLMs. So the style is slightly different. They are maybe a little bit more fluent. And this also affected all of the other parts of how to train them, how to evaluate them and everything. So I think the main progress in MT was mostly into evaluation or from my perspective. And how to actually find out that they are better than previous models, that they are improving. Or it's not just a mirage that humans are biased by the fluency, for example. I think it's well known kind of like the generational kind of milestones of the technology, you know, and I was deeply involved in both the statistical machine translation age, where we, you know, we built dedicated models for enterprise customers.
8:32And then the shift, of course, to the neural age around 2015, and then the advent of transformers. So from the early neural to late neural and then to LLM. But I think the change from neural to LLMs, maybe the most important aspect of that is, first of all, that grammaticality and fluency in most of the major target languages is largely solved. Basically, the problem with machine translation today is not, you know, before, even in the neural age, but certainly, you know, around 10 years ago, if you were using machine translation, you know, you were getting, you know, rough, rough language out there.
9:18there were grammatical mistakes, you know, in actresses related to actual semantics, but also just the generated language was completely rough. And so editing of machine translation 10 years ago was a very different type of endeavor than what it is. Now, late in neural, you know, things got already much, much more fluent, but fundamentally we're still translating sentence by sentence. So still, you know, when you're with translators that are working or any kind of enterprise project that's working on translating a document or a piece of content or a web page, right, had to really worry about how those sentences fit together.
9:59And there was still a lot of editing work. But now with LLM and the move to LLM doing the translation tasks, they do it in full context. They generate highly grammatical language, right? It's very rare to find actually grammatical error. So they do make mistakes, you know, semantic mistakes here and there. But humans translators make mistakes all the time. The characteristics of the mistakes that LLMs make are not necessarily exactly the same. But the fact that they're actually so fluent and can generate so much in context, right? And either you can do that nowadays just off the bat and use a model that is going to generate that type of very fluent contextual language.
10:45Or maybe for whatever reason, there are domains, certainly on a commercial space where people are still using neural models, but then you can use an LLM to basically post-check itself, right, and automatically smooth out and correct and, you know, install things that only humans were able to do before. So in my mind, those are the main really radical differences. And it has deep impact on the workflows that are used across the industry. And we're in a transition period of moving from well-established, well-understood, both in terms of process, but also in terms of economics models of MTPE and post-editing and automated workflows, to LLMs and agents doing various parts of the puzzle.
11:39And that transition is quite difficult and challenging to a lot of users and enterprises that use this technology at scale. I'd like to ask Tom and Alan to tell us a bit more about WMT, what WMT actually is, and why does it matter for the field? Let me start about the general WMT and what it is. So why is it important is that over the last 20 years, every year, we are evaluating what's the state of the art and all the academics and industry can compete with their best machine translation systems and see how far we can actually push the frontier of the research to see how good the systems are. And as well, we are also pushing how to evaluate them.
12:29And every year, we are improving on the pipeline, because if you would be evaluating them the same way like 10 years ago, we could kind of claim that machine translation is solved and there is no need to focus on it. So the main important is that every year, it's completely new set up, a fresher data set for evolution, evaluated by humans. And it's kind of like a competition with among the academics and industry in this realm. And as for what is WMT? So WMT is a conference which also has multiple shared tasks that focus on different subparts of the whole machine translation field. So from a low-resource language, terminology-focused translation, post-editing, etc.
13:12And then there are two main tasks, I would say. The one which is the oldest, it used to be news translation. Nowadays, it's a general machine translation. But the main focus of the task is to train a general purpose machine translation that can translate any domain, any text into a subset of the languages, which is selected every year differently. And then the other most important task is the metrics share task or automatic evaluation share task, which is focusing how can we design automatic metrics that will correctly evaluate how systems are evaluated. Maybe with this, I would give a floor to Alan.
13:58So automatic evaluation of translation has become very, very important because human evaluations are just so slow, expensive and difficult to run. And there's been many evolving use cases, but two main ones is just for machine translation developers is to be able to drive their research, to know when they are actually training a new model or new variants of a model, whether that new model is better than the previous one. You need an automated evaluation system that would actually assign scores and give you that type of information at a granular level as much as possible. And then there's what's become known as QE, which is basically the translation time quality assessment of machine translation or a translation, where the goal is actually to know at translation time what the quality of any generated translation is so that you can actually be actionable, you know, do things, make decisions on that, whether they're workload decisions, whether it's post-editing decisions where editors should focus more attention, what can be pushed out maybe in raw machine translation form and that nature.
15:12So those two scenarios, the main scenarios of kind of offline contrastive evaluation for machine translation development, which is known as reference-based evaluation, also typically called as an empty metrics, that started first around 2002. And then in WMT, starting around 2008, there was a shared task on that. And then QE metrics, which kind of came later, or QE systems, was a separate shared task at WMT, even though there's a lot of similarities. They're both automatic quality evaluation of machine translation. But the two scenarios were different enough and the technologies were different enough that there were two kind of parallel shared tasks for quite a few years.
15:59And this is the first year, actually, where we unified and merged both of those things together because the underlying technology, particularly with LLMs, for solving, for building systems that solve this task have come together very closely together. So this year we did a unified share task on both reference-based and reference without references, so both metric use cases and QE use cases together. And we adopted fundamentally all of the language pairs and the data set up from the translation test that Tom just described to you a few minutes ago. So 16 language pairs, including low resource language pairs, much more difficult data for the machine translation system, generating human.
16:46So here in these tasks, we do generate human assessments because that's the gold standard against which we evaluate all these automated translation quality systems. How well they're doing is evaluated against human annotations of translation quality. And predominantly, this has been over the last several years done using MQM annotation. The listeners probably know about MQM, but if not, we can talk a lot, clarify that. And this year, it was mostly done with a simplified MQM schema that Tom and several other researchers actually proposed a couple of years ago called ESA. It's basically a simplified version of MQM that allows for a little bit more faster and easier annotation by humans and hopefully a little bit more consistency between the annotators.
17:41That was one of the major changes this year, Alan, correct? Yes. So there were several significant changes this year, one of which was the unification of the reference metrics plus NACUE under one unified shared task. So references are optional and the systems that are submitted can use them when they're there or not use them when they're not there. One was the data setup that we adopted from the general shared task, which much larger, longer segments. So it's paragraph-long segments and harder data for the machine translation system. And the third was the human evaluation that's done primarily with ESA.
18:25So, Alan, you mentioned that the test sets used from the NDPs or the tasks were actually harder this year. And Tom, would you like to explain more on the more difficult tests that you use for these tasks? I think the one thing that we've been seeing over the last two years or three years is that there have been more and more claims that machine translation is soft and then maybe there is no need to push more in that realm and like let's focus on something else. But the actual reason that we as a researcher saw is that, well, we've been evaluating on relatively easy sentences and especially for the high resource languages.
19:01translating news article or Wikipedia pages is really easy. That's simple language. And you don't need a specialized system to do that. And with LLMs, they actually do it perfectly. So what we focus this year the most is how to increase the difficulty, how to focus on something that's actually challenging, not just taking random news articles, but instead we focus, for example, on news commentary, some reviews that are more creative language. We focus on the literary domain, which was one long story. We focus on speech translation. We took videos and transcribed them with automatic speech recognition.
19:40So the systems would either have to use the speech or the audio input or know how to fix the errors that are from the automatic speech recognition. And lastly, we focus also on the social domain, basically user-generated content that contains errors, contains different vocabulary and style. So those being like one of the ways how we try to improve the difficulty. Second is that we also developed a difficulty sampling algorithm, which when you have a large pile of data, let's say that you collect news articles, how to select those that will be more challenging for translation. And for that, we developed the algorithm that can sort them, and basically, we took the hardest ones.
20:23So once we had all of these more difficult sources, then we asked a human translator to translate them. And, for example, from some of their feedback, we found out that it's difficult even for humans. We've been getting feedback back that, oh, I don't need more time to translate it. This is more challenging. I don't understand this. Do you have some terminology to support it? So that was one of the first feedbacks that we got. And later, when we also got the automated systems, we saw that they do struggle more in contrast to last year. So they contain much more errors in the output. So it's the direction that we went, basically.
21:04Let's actually focus on something that's challenging and needs more effort to translate rather than simple news articles, for example. So I want to point out what the consequences of that are on actually on the evaluation system shared task. So the fact that the content is harder for the machine translation task means also that it's more challenging for human evaluators to evaluate the quality on that content. So our human annotations for gold standard that were developed also for the main shared tasks, but also for our shared tasks, were more challenging for the annotators to do. On the other hand, the fact that it's harder for machine translation doesn't necessarily mean that it's harder for automatic metrics to distinguish between those systems.
21:57Because, you know, when machine translation was really bad, it was easy to develop actually metrics that distinguish between them. because it was very easy to tell between good, medium, and bad machine translation. But nowadays, when most of the translation generated by machine translation systems are actually very high quality, and indeed, it's not the case for every of the language pair that we saw this year, especially with the techniques that Tom mentioned for making the data harder. It's not that everything was still very easy, but for certain language pairs, we saw it was still pretty easy for some of the content types.
22:37So now it's actually much harder for automated systems to make, to identify, you know, the small differences between those systems, especially on a segment by segment basis, right? So when we run our task, actually, we have kind of two meta evaluation processes there. I don't want to go into too much technical detail, but one of them looks at what we call the system level, which is fundamentally indeed still the ability to distinguish between the quality of two systems on a collection of data that we have, on a collection of documents. And there, because it's aggregated over a collection of data, overall kind of the strengths and weaknesses kind of pop out and automated metrics can do a fairly good job of distinguishing between the systems in the same way that humans do.
23:34But on a segment by segment basis, even with these long segments, it turns out to be quite challenging, both for humans, but even more so for automated systems. And we see that particularly this year. And this year, one of the things is that we tried out, we didn't just collect submissions from participants from the outside, but we, the organizers, not only ran some legacy metrics, basically, to have as baselines and comparisons, but we ran all of the major LLMs with simple prompting as an LLM as a judge for the first time. This is the first time that we actually, in this year, TAS did this. We ran all the major LLMs with promptings to do machine translation quality scoring, and we evaluated how well they do.
24:26And we got some interesting results. We can talk about those in a minute in more detail. So, Mylon, would you like to tell us a bit more about the findings this year? I think that the most important thing or the most important finding this year is how strong the reasoning based large language models were as evaluators, correct? Yes. So, for our shared task for the automated translation quality evaluation, I think a few things kind of popped out and there were also some things that we didn't quite expect. But the main kind of headline, I would say, is indeed we had two very strong LLMs. Both of them were kind of the latest generation with reasoning that did very well in the evaluation, both at the system level, but also at the segment level.
25:17I can mention them by name I don't know if that's important the findings paper has all the details and you can actually see exactly who those are one of those was one of the LLMs that we ran and you know it was fairly basic prompting the other one was a system that was submitted by an outside participant where they not only use a strong LLM with reasoning but they did extensive prompting and setup of the task for the LLM to including very specific examples about how to actually score translations, including detailed instructions about what to generate, etc. And that seemed to be that seemed to work very, very well.
26:12But probably one, we only had very few submissions that actually, to my surprise, actually, I was hoping we would get more submissions of this nature, but we didn't get too many submissions that had detailed prompting with an LLM to do this task. And I think that's clearly in my mind where the future is going. So I think, yes, that's probably the highlight of the shared task this year. One of the surprises was actually that we got a lot of submissions that were variants of trained neural metrics that have been very strong performers in recent years, particularly Comet, but not just Comet. as metric X in quite a few submissions did some, you know, basically took the baseline comment and did something, either extended it to multiple sentences or incorporated other types of knowledge or various things like that.
27:11And all of those variants of those trained neural metrics did not do very well this year. They were surprisingly weak. And we're still in the process of doing some in-depth analysis to try and fully understand really why those metrics underperformed this year. I have a feeling it's mostly, and we have some smoking guns. Tom was helpful also already on trying to find one highlight there. I'll mention it, actually. Maybe, Tom, you can articulate even more about this. But it turned out that many of the submitted machine translation systems that were submitted to Tom's task, to the general translation task, were actually tuned on Comet.
27:55Basically, they used Comet in the process of developing the variant of their model to hill climb or to improve results. And you get better Comet scores as you do that. But any metric has some weaknesses and biases. And if you just hill climb on a metric, then you're incorporating those biases in the output in ways that are subtle and you're not aware of. And then when you evaluate that model and compare it to independent human judgments again, then those biases pop out. And then therefore that model will actually underperform. You know, it will get high comet scores, but those comet scores will not agree very well with human judgments anymore.
28:40So we think that that was one of the things that happened this year, that because many of the systems hill climbed on these metrics, then it caused the metrics actually to underperform on those systems when data from the system was used for the evaluation. But the other thing is just the radically different data scenario. These really long segments, these low resource languages, the hard data, it might indicate that these trained neural metrics were over-trained for data scenarios that were basically no longer, did not match the data that we had this year. I remember you, Alan, mentioning in a call we fired a couple of days ago that a comet became a victim of its own success.
29:27Exactly. But indeed, that is the case. I think the fact that many people in the community adopted Comet as the way to do automated evaluation over the last few years and then used it to train their system, in fact, it is a victim of its own success because now it is underperforming on evaluating systems that went ahead and did that.
29:53Florian:Tom, do you have anything to add to the eval part? I wanted to just highlight why I think it's critical that evaluation is staying on the tiptoes and keep pushing. And it's more important than even the modeling. That basically, this is what we are currently seeing in the field. It's like the classical good heart's law problem. That basically, the moment that the metric became a target, it ceased to be a good metric. And since everyone starts using Comet as a reverse model, for example, or filtering, then the metric can't even catch up. And it's like just one of the puzzles and like why we need to, why it's super critical that WMT is happening every year with new setup, we are updating all the problems that we found in the past and this is just underlying the importance.
30:37Florian:All right, let's go to the performance, right? I mean, we would probably have to screen share because you have so many kind of rankings and tables and it's always challenging for our production team to reproduce them in our tool, but I guess one of the takeaways is that the general purpose LLMs like Gemini 2.0 Pro and GPT, I think you guys used 4.1, did really, really, really well. So what does that tell us, I mean, compared to some of the other systems like things like Google Translate, et cetera, like what does that tell us about LLM-based translation, where it is today? So I would say one comparison is basically like you mentioned a couple of production systems like Google Translate, et cetera, versus the Gemini and GPT.
Read the full transcript
31:22Those are state-of-the-art models that are gigantic, really expensive to run. And of course, if they are best in the class, then that makes them, they've been winning. For example, with Gemini, that was the only model that we run with reasoning on. So that made it like eight times more expensive. And the reasoning actually improved the performance. While in contrast, for example, some of the other models that we would expect to be on top of the class, those are oftentimes focused on the inference speed. And you can't wait like a minute to get a translation as you get with some of the big models.
32:01So that's on the one side, but maybe I would add a little bit more about what I think is really interesting what we saw this year is that how small models built by universities and basically 9 billion models SketchUp with these top systems. And they are actually ranking among the top clusters. It shows that you don't need a 1 billion parameter model. You can actually do it with, well, sorry. I meant for like, trillion. Trillions. Trillions now, man. Yeah, yeah. But what I wanted to say is that instead of having trillion parameter models, you can get to 9 billion parameter models. And they actually rank in top.
32:43So that opens a lot of room for researchers, especially in the academia and in a smaller industry where you can't train for thousands of hours on a gigantic cluster of GPUs. You can train in a couple of days on a relatively small cluster and get actually close to the state-of-the-art performance. Maybe first of all, just to add some thought to what Tom said. I mean, fundamentally, nobody trains models from scratch anymore, right? We're all built on pre-trained models. And the question is, like, what is the starting point and what is done on top of that and for what purposes? And I think what's true is we're going to see that in order to get very good translations for very specific use cases, domains, type of language and things like that, you can do that even in academic institutions and things like that by specializing models with a strong starting point.
33:43And the starting point, you know, that the big guys worry about, you know, pushing the envelope with the latest models and the most general models and the most multilingual performing models. And then other entities, the ecosystem is kind of changing. What you can do with, you know, academic institutions, smaller companies that specialize in certain things is kind of changing. The other thing I think is still kind of interesting is the multilinguality aspect of this. You know, for the very large, you know, economically viable language pairs where there's a lot of data, you know, MT has really gotten really, really good.
34:24But these LLMs in particular, they, a lot of them break down multilingually in performance, including in translation and not just in translation. When you move to smaller, second tier and smaller languages. And part of the interesting thing is that a lot of these pre-trained LLMs, they're trained on the entire internet. And of course, they consume data in all of these languages wherever it exists. But in many cases, they're not specifically dedicated, built in order to truly be multilingual from the get-go. It's kind of almost like an after fact or kind of a nice surprise that they actually are capable of translating and that they are capable, their multilingual capabilities are what they are.
35:14But there's a lot of interesting work on actually making them much, much stronger multilingual, especially for lower resource languages. And then on the evaluation side, again, I think we see sort of the artifacts of that. We see that, indeed, you will be able to get very detailed actual error annotation, automated error annotation, equivalent to ESA or to MQM with very detailed prompted reasoning models. But they are going to be expensive to run. So the question is then how do you actually then distill that or create models that are much more scalable, you know, kind of leaner, faster and more economical to run.
36:07And we're in a transition from this earlier generation of neural models that do QE, that all the commercial providers are fundamentally have models like that that generate scores. but those scores are not quite as accurate, I think, already to the scores that you can get. Well, the evidence shows that they're not quite as accurate as what you can get from LLMs. And so there's going to be indeed exactly what Tom was saying, that trade-off between slower, expensive models that are very accurate in their quality performance versus scalable, faster, and cheaper models to run. and that transition is going to happen, I think, over the next few years.
36:54Very interesting. I would like to go back to Tom about another interesting finding from the MT search task. We saw that human translation didn't land on the top cluster, Tom. How should we interpret the results? I mean, we might have some headlines about human parity. That's a great question, and I'm always pushing back with any findings kind of with WMT, because there's way more open question that we got with this than the other answers. And instead of saying that we cross the human parity, I think the issue is that the translators that are translating out the set are humans. And we are using professional agencies to translate these sets, but yet those are translators doing 40 hours a week translation.
37:42So they get tired, they are pushed by the deadlines, but they have some limited time. So that's one way which reduces the quality. Second, it could be that basically we are kind of forbidding them to use any pre-translation. So they can't do post-editing, they need to do it from scratch, which could affect the way how they work because nowadays most of the translators don't translate from scratch. So there is more of these open problems on how if there are errors in the translation. And then there is also the human evaluation where we kind of see that the human annotators pressure the outputs of LLMs rather than human translation, despite if they need to assign the errors.
38:22The human translation usually has less errors, but if they are assigning which way they like more, they go for the LLM. Maybe because it's more fluent, maybe because it's smoother. That's... we don't know, but those are these various stuff that we see from it. And so I would be really careful about claiming human parity. I don't think we have human parity. maybe we at some point will reach some like human parity on a tired translator doing it 40 hours a week and evening on Friday but it's more about like how can we make it maybe next year better and improve the setup for the humans so they actually can do better job or maybe we can allow them to use the post editing maybe we can actually force the human annotators to mark the errors and if they don't mark error they are not allowed to give lower score so something like that.
39:13I think that The human translation and the human, you know, particularly the human evaluation, it becomes sort of the Achilles heel or the weakest kind of the weakest kind of thing in this entire setup. I mean, we're getting to a point where it's just getting very difficult to get really, really good, reliable, consistent human translations and human annotations with which we can compare basically what the automated systems are doing. Because if we don't actually have a really good, strong gold standard to compare with, we're basically just shooting in the dark or kind of swimming in a fog, right?
39:53We cannot assign really good assessment to what machine translation systems are doing without knowing basically what is right and what is wrong. And they're all getting so good at this point where it's getting very difficult to get very accurate human annotations and human translations against which we can compare and trust that that comparison is actually meaningful. so there's one of the i think the most difficult challenges that i think tom is thinking about this too but you know we're we're really thinking about is how how do we actually move forward in this world in which it's so difficult to do to generate gold standards against which to compare with you know how do we improve human uh translate you know do we do we just get rid of of human reference translations or human you know to compare against uh what would what would replace that do How do we get human annotations of detecting errors to be consistent and reliable and high coverage in terms of not missing errors, right?
41:02When there's so few errors, it's so easy to miss an error here and there. And the annotators do that all the time. They don't so much disagree about errors. It's not that annotator one says, this is the error. No, the annotator says, no, no, no, it's a completely different error. That doesn't happen. but annotator one misses some fraction of the errors and annotator two misses some fraction of the errors and they each miss some different errors. Exactly. Very important. Very interesting. Let's now move to the next question about the actual deployment. And this question can go to both of you. So for teams deploying AI translation and evaluation in VR workflows, what should they take away from these results?
41:43What actually matters in practice? Evaluation methodology is changing, and it is actually a challenge for a lot of, particularly practitioners and enterprises and providers to actually adopt methodology and automated processes that know how to actually leverage these systems that we basically find are performing well into repeatable processes. I think one of the challenges there is that, again, in practice, there are different use cases and applications for things like machine translation, and automated transaction quality evaluation. If you're doing QE, for example, as part of a workflow in order to decide basically how to route content, then you need an evaluation that basically that will look at the system that's generating QE score then see its effectiveness in doing that kind of routing that you want to do.
42:52So for example, maybe you want to just identify that the documents or the pieces of content that meet a certain quality threshold and send those out without any human review or without any human editing. So create a testbed scenario that looks at how well a QE system or an LLM as a judge is capable at doing that triage of the highest quality versus all the rest. Or maybe you have a limited budget and what you're most interested in is putting your post editors on the lowest quality translated documents. So in that scenario, it's a different triage. It's identifying the lowest quality reliably. So you have to set up actually a workflow or a scenario that will evaluate the specific automated translation quality QE system, regardless of who the provider is, at its ability to detect the lowest quality documents or the lowest quality segments, right?
43:58And we don't actually have a published blueprint that is out there for everyone to follow and for this, an exact recipe to follow in evaluating different technologies or different systems available on the market for these different types of use cases. And I think that's actually a gap that I'm looking to try and address. We're going to try potentially forming a working group across the industry that is interested in these types of questions and maybe can come up with some guidelines or recipes about how to evaluate the performance of automated translation quality systems for these specific commonly used use cases.
44:51Florian:We could hear, I mean, you guys launched Command Translate so recently, which I think does, the general family of command would address the multilinguality from kind of inception that Alan before said is a bit of a problem. But okay, going back to the original question. So your takeaways for maybe here from something like the W or from WMT, like as a practitioner, are you bringing this back? Are you talking to the, to proglet people there, to other researchers? How does that work? That's an excellent question. I would say that maybe I'll answer a little bit from like the modeling side of the house rather than the evolution.
45:35So one of the points is basically like how to improve on them, how to know that we are improving and what to do. So unfortunately, like, so what we saw from WMT is that, yeah, rework modeling, filtering with Comet works. There is MBI. There's plenty of techniques that actually works on the modeling to push the performance. But the key point is more about what to trust. As Alan mentioned, so for example, the human evolution are still gold standard. So it's much more difficult to bias humans to judge the systems. So maybe my main answer would be that if you're a practitioner in a company and you do not strain your own models, you need to be really vigilant about how to evaluate them.
46:25So if you don't know how it was strained, it can be biased towards your evaluation. However, if you build your own models, then the key takeaway is, for example, don't use the same metrics that we use for the, like, if you use commands somewhere in the pipeline, do not evaluate on it and maybe go with LLM-based metric as LLM as a judge. and basically what I think of is the main feedback is more about like convincing for example if we would be selling the command a translate is like getting customers on the side that we need to evaluate on difficult test sets so for example not evaluating on Flores on German just because for those who don't know Flores is Wikipedia articles it's amazing resource for low resource languages but most of the people and papers are using it for German, French, Spanish, Italian, Chinese.
47:20For these languages it's overfitted and there is no progress and then for example if you would have a customer who came and like oh yeah on the Flores and on Blascore you are this bad and that then WMT uses a could be used as a platform to show like yeah maybe that's not ideal evaluation Let's compare on something reliable or more difficult or more advanced.
47:43Florian:Makes sense. Makes sense. I want to close on a couple of big questions or a series of big questions. Alan, let's say Safaba, 2026, you're starting again. What's the frontier? What's the current frontier in AI translation? Like what areas for growth are there for anyone offering it at this point in time or maybe wanting to start? Like what are the big remaining unsolved problems? And I said it was a series. Why do they remain unsolved three years into the LLM revolutions and, you know, like hundreds of billions and soon trillions of dollars spent on this and nuclear reactors powering all of this?
48:22Okay. Those are hard questions. But they're two separate ones.
48:25Florian:Like if I were to start a Sofaba today or what is like something that is interesting? So I think personalization. I mean, we already know that it's becoming very much, much easier to actually adapt translations for specific scenarios by just by providing the context and providing detailed information. You can basically bias the, you know, or instruct the LLM basically to generate translations that are very specific. And I think there's lots of opportunity for personalization, particularly like in marketing situations where you're tailoring to specific communities, to specific groups, but all the way to even personal personalization.
49:06So you can have a situation in which, you know, instead of companies generating, just translating, you know, a particular content, say marketing content into a target language in which, you know, the company is operating, it actually generates on the fly a translation for you, knowing everything that it knows about you. so that that translation is going to be much more impactful, and particularly in situations where they're trying to persuade you to buy something, for example, right? It will have immense impact on persuasion and things of that nature. So I would go, I think there's an opportunity there for technology in that very, very detailed level of personalization.
49:53But unsolved big problems. I think the ability of these systems to still really self-reflect the way that humans do and check themselves. Like, why are we still in a world in which the LLM that's generating the translation doesn't already know everything that there is to know about the quality and doesn't generate errors to begin with, right? Why are we still in a world in which we have to have a separate system for measuring translation quality? And the translation system itself is not self-reflective or capable of avoiding making those errors or knowing how confident it is exactly about what it's generating.
50:40And those are fundamental questions about the way that these neural systems fundamentally operate, right? And do they have enough self-consciousness about what they're doing? And I think the answers to those things are evolving. We don't quite understand them yet. But there are fundamental things about the way these models operate that we still don't quite understand, and therefore there are fundamental things that we haven't been able to solve yet.
51:10Florian:Throwing them a bunch of harder problems like what you're doing, Tom, is probably going to be a lot more helpful, right? Than, you know, trying the same thing again and getting these scores, getting super high scores, right? I mean, part is to give the systems a ton more challenging tasks to solve. Anything you want to add here, Tom, to close this out on current frontier, unsolved problems, thoughts about the future? I want to summarize it nicely. My first thought would be some personalized translation. So just pushing the state of the art is not the frontier that I would focus with if I would start a new company.
51:50I would focus on the personalization, basically adding more context. Like if you have the marketing campaign, here's the context, here's the target country, do translation that will be targeted for this. It's like focusing on models that can do more than translation, being more multilingual, focusing on other tasks. Secondly, maybe something new to add to the discussion is I would focus more on agenting and combining multiple systems and maybe not focusing on how to make one system the best in everything, but actually, I can't relate to a long process, so that having one system that I can rely on being the best in quality estimation, the second is amazing in translation of the terminology.
52:32Other one would be perfect in like Slavic languages. So basically combining them together and knowing which one are related or what, putting them together could be actually a way how to improve on all of the fronts. I don't disagree fundamentally. I think the two things are complementary but specialization is clearly much, much easier with agentic systems. Specialized agents interacting with each other, very powerful, no doubt.
53:04Florian:Maybe just one last thing there. The agentic thing kind of puzzles me sometimes because what if one part breaks? Don't you add a lot more potential for error? Like if one gets it wrong, one of these agents, you have like a failure cascade throughout all the other agents? Not necessarily. No? Okay. They can self-correct. Well, yeah, or another agent can basically check and verify the work of the first agent. So when the first agent makes a mistake, there's going to be another agent that has a different perspective and specialization that will catch the error and flag it or correct it, etc. Awesome.
53:41Florian:This was great. Thank you so much. Thanks so much, Alan. Thanks so much, Tom. Thanks so much, Maria. Thanks for taking the time today and thanks for all the great work you do with WMT. Super, super valuable. Thank you.
From the publisher
Tom Kocmi, Researcher at Cohere, and Alon Lavie, Distinguished Career Professor at Carnegie Mellon University, join Florian and Slator language AI Research Analyst, Maria Stasimioti, on SlatorPod to talk about the state-of-the-art in AI translation and what the latest WMT25 results reveal about progress and remaining challenges.
Tom outlines how the WMT conference has become a crucial annual benchmark for assessing AI translation quality and ensuring systems are tested on fresh, demanding datasets. He notes that systems now face literary text, social-media language, ASR-noisy speech transcripts, and data selected through a difficulty-sampling algorithm. He stresses that these harder inputs expose far more system weaknesses than in previous years.
He adds that human translators also struggle as they face fatigue, time pressure, and constraints such as not being allowed to post-edit. He emphasizes that human parity claims are unreliable and highlights the need for improved human evaluation design.
Alon underscores that harder test data also challenges evaluators. He explains that segment-level scoring is now more difficult, and even human evaluators miss different subsets of errors. He highlights that automated metrics built on earlier-era training data underperformed, particularly COMET, because they absorbed their own biases.
He reports that the strongest performers in the evaluation task were reasoning-capable large language models (LLMs), either lightly prompted or submitted with elaborate evaluation-specific prompting. He notes that while these LLM-as-judge setups outperformed traditional neural metrics overall, their segment-level performance varied.
Tom points out that the translation task also revealed notable progress from smaller academic models around 9B parameters, some ranking near trillion-parameter frontier models. He sees this as a sign that competitive research is still widely accessible.
The duo concludes that they must carefully choose evaluation methods, avoid assessing models with the same metric used during training, and adopt LLM-based judging for more reliable assessments.




