In short
SlaterPod #290 with Laniqo (Lanico) CTO Artur Nowakowski on building secure, adaptive AI translation. He emphasizes trust, control, and consistency; using open-source models that can run fully private/on-prem; domain adaptation via glossaries, translation memories, style guides, and layout preservation for documents (especially PDFs).
Key claims
LLMs are strong for generic translation, but domain fine-tuning and model swap-ability matter; open-source supports EU privacy/sovereignty needs; at e-commerce scale, cost and error control require MT quality estimation/assurance workflows.
Notable examples
WMT 2022 shared task win (Czech↔Ukrainian) leading to company founding; Allegro (Central/Eastern Europe e-commerce) using custom domain models for hundreds of millions of offers; “Format” dataset (4,000 PDFs, 15 language pairs) on reconstructing PDFs with OCR + vision-language models; CompactQE highlighting MQM error spans with proposed corrections.
Guests
Artur Nowakowski, co-founder and CTO of Laniqo; host Florian SlaterPod (interviewer).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOArtur Nowakowski's Background
0:45 to 2:15
Discussion about Artur's academic background and the founding of Lanico.
“It's quite hot today, but it shouldn't be an issue today.”
Lanico's Origins and Growth
2:15 to 3:45
Exploring the journey from academic research to founding Lanico as a company.
“when I was doing my PhD I did the century of Artificial Intelligence was founded at the AMU.”
Understanding Lanico's Mission
3:45 to 5:15
Artur discusses the key problems Lanico aims to solve for businesses.
“So this basically, our win and the share task turned into a larger industrial endeavor.”
Navigating the Post-LLM Era
5:15 to 6:45
Insights into how Lanico adapted to advancements in translation technology.
“It's trust, control, and consistency of the translations.”
Open Source and AI Translation
6:45 to 8:15
Artur shares his thoughts on the importance of open source in AI translation.
“Now, why don't we dwell a little bit on this kind of the 2022, 2023 moment, right?”
Language Coverage Challenges
8:15 to 9:45
Discussion on the quality of translation models for different language pairs.
“So in our case, we started, of course, with NMT models that we built ourselves and embedded them into the product.”
Allegro: A Key Customer
9:45 to 11:15
Artur introduces Allegro and its role in AI translation for e-commerce.
“But like in our case, when we approach this problem, we usually experiment with multiple research models, including Chinese ones, European ones, US ones, and then we try to find the best that fits to our problem.”
E-commerce Translation Challenges
11:15 to 14:00
Exploring the challenges of translating for a vast e-commerce catalog.
“to keep this more private and more laser focused to specific domains, to specific customers.”
Scaling Translation for E-commerce
14:00 to 16:32
Learn how custom models solve translation challenges for large e-commerce platforms.
“and we've been helping them with this challenge of expanding to new countries.”
Quality Assurance in AI Translation
16:32 to 18:24
Discover methods for ensuring accurate translations at scale in AI systems.
“or is it kind of ticking over quite nicely?”
Show all 17 chapters
Security Concerns in AI Translation
18:24 to 20:34
Understand the security considerations for AI translation in sensitive industries.
“You already get a list of text with highlighted issues, for example.”
Business Development Strategies
20:34 to 22:44
Explore effective strategies for generating leads in the AI translation market.
“What's your approach to business development?”
Innovations in PDF Translation
22:44 to 26:10
Learn about new methods for preserving layout in PDF translations with AI.
“So you recently launched something called Format, and we spoke about it briefly before.”
Language Resource Creation for Clients
26:10 to 28:00
Find out how language resource creation helps clients improve their translation processes.
“Also you have something called language resource creation, right?”
Challenges in Translation Memory Management
28:00 to 31:04
Learn about the challenges companies face with translation memories and how to address them.
“So basically, yeah, I mean, the scenario to me would be marginally sophisticated company that has a certain amount of already kind of translated or content in the past, localized content.”
The Importance of Staying Connected to Research
31:04 to 33:14
Discover how maintaining ties to the research community provides credibility and collaboration opportunities.
“You go to the conferences, you publish papers.”
Future Directions for Lanico's AI Translation
33:14 to 34:23
Explore Lanico's future product priorities and the move towards multimodal translation solutions.
“for the next, I don't know, second half of 2026 and maybe heading into 2027?”
Transcript
Automatic transcript. May contain errors.0:00Right now, I think that the most important things that we are trying to solve is trust, control, and consistency of the translations.
0:12Artur Nowakowski:All right, and welcome to another episode of SlaterPod. Today on the podcast, we welcome Artur Nowakowski. So Artur is the co-founder and CTO of language technology platform, AI translation platform, Lanico. Hi, Artur, and thanks so much for joining today. Hi, Florian. Thank you very much for having me. Glad to be here. Absolutely. This is the first episode in like three weeks or something. So, thanks for breaking the summer break for us here. So, Artur, where are you recording this from today? What country, what city? Currently in Cozumann, in Poland, where I'm recording from our office in the middle of the city.
0:50It's quite hot today, but it shouldn't be an issue today.
0:53Artur Nowakowski:I was complaining about it in the last couple episodes. I'm complaining about it today. We've had this like three-month heatwave over here in Europe. So anyway, cool. So look, let's talk a bit first about your background. You have a very strong background in academia and then kind of journey to co-founding Lanico. I understand that Lanico grew out of the MT research at Adam Mitskiewicz University and so tell me a bit first when you maybe tell us a bit more about your academic background and then like when you realized kind of that research that you were working on, you know, could maybe become a viable commercial product?
1:31Regarding the history of co-founding Blanico, it's quite an unusual and interesting story. I mean, it's probably worth mentioning from the beginning that the Adamiskewicz University has a long history of research in machine translation. Right now, it's about 30 years, starting from real-based machine translation and going through statistical, then neural. also it's worth mentioning that Marian NMT framework that was then later developed at University of Edinburgh and Microsoft was initially developed at our university and then right now still going research is going strong with DLMs so in 2022 when I was doing my PhD I did the century of Artificial Intelligence was founded at the AMU.
2:27And I was the leader of the machine translation team, which was quite small back then. It was only me and a few master students. And we decided to take part in WMT in 2022,
2:51where we took part only in two language directions. It was from Czech to Ukrainian, Ukrainian to Czech. But we managed to win the share task back then, beating all the other constrained and unconstrained solutions. And it turned out to be quite a big achievement that was quite populated in the media. and after some time, Hans Langenzeit from Germany reached out to us saying that, well, they've been looking for a very strong team with a research background and machine translation and that they would like to try to collaborate with us and maybe start a company in the future. So this basically, our win and the share task turned into a larger industrial endeavor.
3:59And a few months later, we founded the company.
4:04Artur Nowakowski:So Pons, that's P-O-N-S for those who don't know, used to be like a, I don't know, way back in it, it was like a dictionary company, right? I mean, when I did my translation study, obviously I had my pawns as kind of a book. So obviously they've evolved very much over time. So they are now like an investor, an owner, is it kind of a joint venture or how does that work? So they are the main investor and an owner of the company since the beginning. And also Adam Itskiewicz University itself is also a shareholder in the company. So we are like a spin-off from both worlds, from the university and from Pons Lannenscheid as the main investor.
4:51And they are still using our products in the German market and acting more as a distributor right now of the technology.
5:05Artur Nowakowski:So maybe let's go to like the, I mean, it's a very young company, right? So what's your, I guess, elevator pitch? Like what problem are you ultimately trying to solve for businesses? Right now, I'd say it's three things. It's trust, control, and consistency of the translations. I mean, right now we are in the post-LLM era. When the company started, when there was an idea for the company, there wasn't just GPT in the middle of 22. but right now I think that the most important things that we are trying to solve is first of all the security of the translation so what we all do is based on open source models that can be fully private can be also deployed on premises we are also trying, we are also solving the adaptations for specific brands like e-commerce, medical or automotive, for example, where we adapt the translations to the specificity of the brand with the glossaries, translation memories, style guides, specific preserving.
6:21When we are translating documents, we also are preserving the layout of some very specific documents. That's not... And I don't think this is a solved problem yet.
6:34Artur Nowakowski:We had somebody on the podcast based in Singapore who basically built a company Razor focused on preserving the documents just for the finance space, right? So yeah, it is a tough one. Now, why don't we dwell a little bit on this kind of the 2022, 2023 moment, right? So you started like literally a few months before ChatGPT. Like how did that feel? Like you started something, you had an idea, the thing came out, we all play with it and kind of the rest is history, right? But you're still here. We're recording this podcast and your company works. So walk us through the different kind of moments and the thinking you had when you realized, okay, there's a kind of a revolution happening here and I'm at the absolute epicenter of it.
7:21Yeah. Yeah, I mean, that was quite a pivotal moment for us in the end of the year, I think, when the LLMs actually started getting better at translation. Because in the beginning, when TGPT came out, still for some period of time, I mean, standard NMT models were still faster and even better in quality. but the longer that time went, we all saw that the elements are just more versatile and can be used for more things than just translation. But from the very beginning, we've been building our product and the whole approach in mind that we should be able to easily swap the main technical things that are sitting inside the product.
8:20So in our case, we started, of course, with NMT models that we built ourselves and embedded them into the product. And then in 23 and 24, we had to switch to LMS because the way we did it is that we used open source elements that we are able to control, we are able to fine-tune, so we are not just taking things off the shelf and embedding them in the application.
8:55Basically that allowed us to make sure that our product is more versatile and that we are not 100 % reliant only on NMT models, because if NMT models would be our main product, then it would be very difficult for a company to stay above the water.
9:17Artur Nowakowski:So you mentioned open source. There's been so much talk recently about open source with the LLMs with all the Chinese models coming out and these various rankings being on par with the closed source ones. What are your thoughts around that and its relevance for AI translation? From what I see even in the benchmarks is that well, right now LLMs, especially commercial ones, are very good at the generic translation quality. It's very, very hard to beat them. But like in our case, when we approach this problem, we usually experiment with multiple research models, including Chinese ones, European ones, US ones, and then we try to find the best that fits to our problem.
10:13And that's also what we describe in our research papers. But right now, I think the ability that, well, you are able to take the open source models but fine-tune it on the data to make it very specific to the problem, to the specificity of the domain is very important. And also regarding the, for example, the sovereignty and the privacy of these models. You are able to host them yourselves in a European data center, which is very important for the customers that are relying on, let's say there is an EU AI Act coming in with the new regulations that are right now and there will be more. And I think the open source models allow us to keep this more private and more laser focused to specific domains, to specific customers.
11:26Artur Nowakowski:What about language coverage? I mean, you mentioned before Czech, Ukrainian. Like on many of the kind of business relevant language combinations, do you feel that, you know, we've reached kind of a plateau, it works well now. Yeah. You basically, you as a provider of kind of at scale machine translation, and we talk about it later for e-commerce, for example, like you're taking these models, you're customizing it and then you launch them. but the underlying models are good at most language combinations now or in some areas they're still lacking? I would say that it heavily depends on the language pair.
12:08For some of the languages that are very high resource, let's say German, English, Spanish, French, the quality is quite high for most of the open source models. For the lower-researched ones, there is much more work to be done still. For example, with the Allegro case, we are more focused on this region-specific languages, such as Polish to Czech, Hungarian, Slovak, Ukrainian. And in these languages, the open source models are not, like, there is still room to improve. And the further we go, like, if we would go to some more rare language combinations, even in African languages, then there is still a very large gap, I think, in both commercial and open source models.
13:16Artur Nowakowski:So you mentioned Allegro. I think they're on your website also as a key customer, right? So for those of our listeners who are not familiar with the company, tell us a bit more about Allegro. And then you mentioned some of the challenges being the language combinations here with languages that are less well covered by AI translation. But what are some of the other challenges? So first quick intro to Allegro and then some of the other challenges for AI translation in e-commerce. So Allegra is the e-commerce company, actually the largest e-commerce company in the Central and Eastern Europe right now.
13:50They originate from Poland, from Poznan. And in the past few years, they've expanded into Czechia, Slovakia and Hungary. and we've been helping them with this challenge of expanding to new countries. Originally, the product catalog we have is hundreds of millions of offers covering multiple domains. They wanted to have very accurate translations that would allow customers in these other markets to be able to buy these products and not lose credibility at the same time. So the translations had to be of the quality that would not make customers not trust the company. And I think there are a few big problems here.
14:56when you are doing this at such a big scale that you have hundreds of millions of offers each day, there are thousands of new coming in that have to be translated. So first of all, I'd say it's a deployment and scalability problem that in this case, if you were to use just big LLM and try to translate everything, then the cost would go into millions of dollars. And that's something you usually want to optimize. So in this case, we are providing custom models for this domain where we focused also on most of the things like terminology because in e-commerce you have, for example, we are translating an offer title, a product title, there is usually, it's usually like four words, five words.
16:01And so you have to be able to also introduce more context into the translation, like instead of relying on just this four or five words. So like introducing that external context, such as more product information, metadata, images, this has also been a very big issue that we had to solve.
16:28Artur Nowakowski:I mean, you solved it, I guess, right now. It's like fully operational been for years, or is it still a big challenge when they roll out new campaigns, or is it kind of ticking over quite nicely? I mean, we've been collaborating with Allegro since 22. I mean, even before the company was founded, we started the collaboration as the university and then continued the company. but I mean the launch today this new architecture has been successful and there are like customers trusting the output of the translation but at the same time we are now experiencing other issues that we are solving for example if you are translating at such a big scale you have to be sure to also limit some critical errors and it's very difficult to limit these errors at such a scale because if for example you have a team of linguists they are not able to verify hundreds of millions of authors so what we are also working right now on is to develop quality estimation, quality assurance workflows that make these things easier.
17:56For example, that you automatically try to detect errors in as many offers or product listings as possible. And then humans can verify them and this also makes their work faster and easier that you do not have to guess where might the error be. You already get a list of text with highlighted issues, for example.
18:31Artur Nowakowski:MTQE, a very tricky one to work on. So when you talk to customers or leads other than Allegro, what are some of the kind of key questions in mid-2026 when you come in with an AI translation pitch? We usually talk to companies that are relying very much on security and privacy right now. So government agencies, government organizations, and medical ones, so they usually want to make sure that what they translate is safe. So this includes, usually the questions are about the security on premise and how we deal with that. Also, like I mentioned in the case of Allegro, there is, like, we usually approach cases that, well, people usually have some translation engine that they've been using for years.
19:35I mean, it's usually quite hard to get into this market and to people that have established workflows and established engines that they've usually been using for at least a few years. So then the question, then there is a much bigger focus on MTQE and automated post-editing side. That, well, they get output from the translation from the translation they may already have but they usually have no idea about the quality of the outputs so then what we can offer them is the is this approach to better detect and highlight errors in such translations so So their post-editing process is much easier. So these are usually the LQA teams in such organizations.
20:38Artur Nowakowski:How do you generate leads? What's your approach to business development? Are you going out there and doing it yourself? Do you have a team? Is it inbound? How does it work? I mean, this is split into a few different categories. So one thing, as I mentioned, we have Pons as a distributor in the dach market, where the brands, the companies over a century old, and the people, they trust the brand and the quality that comes with the brand usually. So, I mean, we have both inbound and outbound channels. And usually, I think inbound is right now more effective. That's usually people that have some real existing issues with the current approaches, their current workflows.
21:35For example, with this MTQE or better adaptation from their specific domain, they come to Poland because they trust the brand, they trust the quality. On the other side, in Poland, we also are approaching this from the other side. We are a university spin-off and Adam Iskiewicz University is one of the top three universities in the country. so this way it's easier for us to approach people that say hello, hi, we are coming from the university we have research background, credibility in the research papers and then it usually makes the conversation much easier that the customers know that they are talking to people that have some scientific background in this field.
22:44Artur Nowakowski:So you recently launched something called Format, and we spoke about it briefly before. So it's Format with a capital F, capital M, capital T. So it's, I guess, would you call it a product or a feature? And then what does it do, like in obviously preserving the format? But yeah, tell us more. Format is actually a data set that we created and we published recently, two months ago at the AMT conference. And it's a result, like it's connected to a feature that we've been developing. Because one of the main features of our product is document translations. And also in this document translation is the PDF translation.
23:30And so we've, like one of the main issues in PDF translation right now, especially if we are talking about scans or PDF, other PDF documents that have no text layer is the difficulty of translating that documents while preserving the layout. So, I mean, you can easily extract the text via OCR methods or some other methods, but actually preserving the layout in the output file for PDFs is a difficult problem. So formats we created while we've been working on the approach for, but much more novel approach for PDF translation. It's basically a data set and a benchmark of around 4 ,000 PDF across 15 language pairs.
24:20And we used it to verify our approach to this problem. So in the standard scenario, like most industrial providers do when you are translating a PDF, you are basically relying on conversion to Microsoft Word or some other format. And that usually tends to break the layout of the document. So what we did instead is that we combine OCR methods with visual, vision language models. and we basically reconstruct the document from scratch without relying on this conversion to Microsoft Office. And this dataset allowed us to properly verify that we are doing this correctly.
25:21Artur Nowakowski:And so authors can use that too, and you said you launched it at, or you introduced it at EMTA? We introduced it at EAMT. I keep getting those mixed up. There's a couple of those. Yeah, EAMT. It's a sister conference of AMTA. But yeah, anyone can use it. It's published on Hugging Face. And yeah, we are looking forward to actually people using it because that's like the proper layer of preservation of PDFs is still a very, it's a big problem in MT. and I mean we are developing a method that can treat this problem better and in the near future we may have new papers with better approach to this issue.
26:18Artur Nowakowski:Also you have something called language resource creation, right? The offering, so like who's using this? you're using eternally? Is there like clients can use it? Partners can use it? Tell us more about that. It's a relatively new feature that we have embedded in our product. It's basically like maybe coming from the problem. Like people, the customers that we approach, especially some smaller customers, they usually have a history of translated documents, but they don't have any style guides, translation memories, glossaries, and they would like to use our product. But the main selling point of our product is the adaptation to customers' language and their specificity.
27:08So this language resources, language resource creation module, is basically a creator of glossaries and translation memories that you are able to upload historical documents in multiple formats in PDFs, Microsoft Office, PowerPoint presentations, etc. And it basically creates these resources for you. And it also filters them via our QE, so that what the customer sees as the output is not just randomly aligned fragments of sentences, but instead already a list of proper terms or segments in translation memory that they can just directly use in their translations.
27:58Artur Nowakowski:Is this mostly like a customer onboarding tool for you? So basically, yeah, I mean, the scenario to me would be marginally sophisticated company that has a certain amount of already kind of translated or content in the past, localized content. They're coming in, they want to work with you. But again, they don't have a sophisticated kind of TMX kind of repository, right? Is that accurate or does anybody else use it? I think there are two main use cases. The first one, as you mentioned, is the onboarding one, that someone doesn't have any resources and wants to use it. And the other is that someone has already translation memories or glossaries, but they are quite noisy, let's say.
28:42We usually have that experience with companies that preserve translation memories, that have been creating translation memories for the past dozen years. These translation memories are then quite big, and at the same time, they are quite noisy. And they can use these automated filtering methods of our QE to clean these resources. Because usually if they would use these resources as they are, there is high chance that they could break something in the translation rather than fixing it.
29:26Artur Nowakowski:You mentioned QE. So is that compact QE or is the compact QE still kind of experimental? Because I read up that you're using kind of smaller open-weight LLMs for compact QE. This is actually CompactQE. This is part of the product that we have embedded in our application. Usually when we come to research and mixing the product with research, it's usually that we have some industry problem that we want to solve. We do some research and we write the results in a research paper. But in the case of CompactQE, we did a lot of experiments with these open source models and we wanted to find a way to approach this problem better so that we do not only give the score from 0 to 100 like with some other methods.
30:25But at the same time, rather than just a score, we highlight the error spans that are in the text, including the MQM categories. And we give the corrections, proposed corrections for each error. And at the same time, we give the whole post-edition proposition that the translator can either approve or not as a correction. So then basically when they go through the document, they already see where the errors are.
Read the full transcript
31:03Artur Nowakowski:So you still have a very strong connection to the research kind of community, if I understand this correctly. You go to the conferences, you publish papers. If I'm running a startup and I want to get traction, I want to talk to clients, like what's your tangible benefit of kind of keeping one foot in that research community? You could also argue like, hey, I've done my part, I've published my papers, I got my PhD, now it's time to build and scale and do all the kind of corporate stuff. Yeah, I think for the company of our size, as we are quite small, like keeping in touch with research and publishing new papers every year, the biggest benefit it gives us, it gives us credibility because we can evaluate our approaches, for example, at WMT, or even when reviewers are reviewing our paper and they accept it for a conference, it gives some sort of confirmation that what we are doing is in the right direction.
32:03And for the company of our size,
32:08we do not have budgets for marketing and sales. The company is 100 times our size. So it acts as a big credibility factor that when we approach customers, that we can highlight that. And also, at the same time, we are collaborating with universities, especially with Adam Iskiewicz University. There is also an industrial PhD program in Poland and that involves a PhD student that's employed 100 % time at the company and there is an agreement between a university and the company. And we have a few such PhD students that they are able to finish their PhD, do their PhD that's related to the work done at the company.
33:07And at the same time, to do this PhD, they have to publish some research papers.
33:13Artur Nowakowski:So to close off, what are some of your priorities for the next, I don't know, second half of 2026 and maybe heading into 2027? Regarding product, I think the biggest priorities would be going more multimobile. Right now we are, like, because right now we've been mostly focused on the text translation, but we see that, like, for example, going more into the voice or images is usually the, could be the right direction to go. And at the same time, we want to explore things in the language AI area that are not strictly related to just machine translation.
34:06So because right now we see that, let's say, even right now and in the next few years, it might be even more difficult. The market for just text translation is quite saturated right now. And with the LLMs, having a very good quality for generic text, going deeper into the domains is, I think, one way and also going more on the privacy. But, well, we see that in the next few years. We should go and extend what we are doing outside of just translation.
34:59Artur Nowakowski:All right. So if people are looking for Lanico, it's written with a Q, so L-A-N-I-Q-O.com. So head over there and check out Lannico. Well, Arthur, thank you so much for taking the time today. This was great.
From the publisher
Artur Nowakowski, Co-founder and CTO of Laniqo, joins SlatorPod to talk about the language technology platform’s (LTP) origins, business model, and research-driven approach to AI translation.
Artur shares that Laniqo emerged from machine translation research at Adam Mickiewicz University after his team won a WMT 2022 shared task, attracting PONS Langenscheidt, which became the company’s main investor and helped commercialize the university’s technology.
Laniqo initially developed its own neural machine translation models but shifted toward open-source large language models as their translation capabilities improved. According to Artur, controlling and adapting these models remains essential for domain-specific use cases and lower-resource language pairs.
He highlights Laniqo’s work with Central and Eastern European ecommerce platform Allegro, where the LTP supports the translation of hundreds of millions of product offers. Key challenges include scalability, cost control, terminology, limited source context, and detecting critical errors across volumes that human linguists cannot review manually.
Laniqo is also developing quality estimation tools that identify error spans, assign MQM categories, and suggest corrections. Its recently published ForMaT dataset supports research into PDF translation that preserves document layouts without relying on conversion to Microsoft Word.
Looking ahead, Artur outlines how Laniqo plans to expand beyond text translation into voice, images, and broader language AI applications, while continuing to prioritize privacy and deeper domain adaptation.




