#287 How the Market for AI Data Has Become a Major Growth Opportunity

12 Jun 2026 · 35 min · 17 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Data-for-AI market growth opportunity; how “deployment data” (data for adapting, aligning, evaluating, and safely deploying AI) has shifted demand from basic labeling to expert-driven alignment, red-teaming, and evaluation.

Guests

Anna Windham, head of research (career began at Appen working on speech-to-text data labeling; led a new data-for-AI market report). Host Florian (SlaterPod) interviews her.

Key claims

Market is ~9.3B (external commercial spend only; excludes frontier labs’ internal work). Demand moved downstream after the 2022 LLM surge toward making models reliable, safe, controllable, and usable in real business workflows. “Deployment data” is iterative and requires domain experts to define “what good looks like,” run adversarial tests, and provide evidence for deployment.

Notable examples

Physician selecting safer medical responses; OpenAI GPT-4 red team with 100 external red teamers across cybersecurity, healthcare, law, finance, psychology, misinformation, child safety, linguistics, 45 languages, 29 countries; frontier labs hiring investment bankers to create evaluation frameworks; Data Mundi folding napkins 100 times/day to teach models.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding the Data for AI Ecosystem

0:45 to 1:55

Exploring the companies involved in the AI data market and its evolution.

“So let me just give you a quick run through kind of the high level bullet points before I defer to you on the various aspects that we covered in this report.”

Anna Windham's Background and Insights

1:55 to 4:10

Discussion about Anna's expertise and relevance to the AI data report.

“And maybe Anna, before we start with kind of the core content of the report, tell us a bit more about your personal background and how this is relevant for this particular piece of research.”

Key Changes in the AI Data Market

4:10 to 6:25

Examining shifts in demand and the role of domain experts in the AI data market.

“That's no longer the case because while they had to scale back on some fronts, others kind of exploded.”

Misconceptions About Data for AI

6:25 to 9:05

Addressing common misunderstandings in the AI data landscape.

“So instead of relying on those large pools of generalist annotators, there's now a new and increasing need for domain experts who can provide judgment that will shape model behavior.”

Market Size and Composition

9:05 to 11:24

Analyzing the estimated $9.3 billion market for AI data and what it encompasses.

“Obviously, people are very interested in that.”

Research Methodology for AI Data Report

11:24 to 13:40

Overview of the research process used to compile the AI data market report.

“I guess one there would come to mind is we had Casper Grothwell on the podcast from Oxford.”

The Rise of Deployment Data

13:40 to 14:00

Introducing the concept of deployment data and its significance in AI model adaptation.

“So thank you to everyone who spoke to us.”

Understanding Deployment Data

14:00 to 16:40

Learn about the concept of deployment data and its critical role in AI model adaptation.

“And they really validated what we were hearing from the providers, which is we need these human experts to take part in these alignment adversarial testing and evaluation processes.”

The Role of Alignment and Red Teaming

16:40 to 19:40

Explore how alignment processes shape AI model behavior and the significance of red teaming in identifying vulnerabilities.

“Let's talk about, you mentioned that red teaming and also the alignment.”

Importance of Evaluation in AI

19:40 to 23:00

Discover the critical nature of evaluation in AI model release and the role of domain experts.

“and then to detect when they've been bypassed.”
Show all 17 chapters

Shifts in the AI Workforce

23:00 to 28:00

Examine the evolving workforce dynamics behind AI development, emphasizing the need for domain experts.

“which is that the more capable the models become the more important it is that we have humans who can tell us in what ways they're actually useful, if they are trustworthy, if they are safe.”

Current AI Data Buyers

28:00 to 28:35

Learn about the major buyers in the AI data market and their spending habits.

“And then you have, you know, Mistral and Cohere, et cetera.”

Emerging Demand: Sovereign AI Programs

28:35 to 29:23

Discover the rise of national AI initiatives and their impact on data needs.

“But apart from that, we have enterprises and AI product builders.”

Diverse Buyer Needs and Evaluation

29:23 to 29:53

Understand the different needs of various AI buyers and how they evaluate suppliers.

“So they need, obviously, language coverage and they have specific security requirements.”

The Role of Language Solutions Integrators (LSIs)

29:53 to 30:50

Learn about the advantages LSIs have when entering the AI data market.

“There's some amazing tables in there that if I was a salesperson in this industry, I would want to get my hands on.”

Transformation in Language Services

30:50 to 32:52

Explore the transformation needed for LSIs to succeed in AI data-related roles.

“So on the one hand, the assets or the strengths that LSI's have are the multilingual contributor network, experience managing distributed workforces, of course, expertise in language evaluation.”

AI Data Industry's Role in AI Infrastructure

32:52 to 34:33

Examine how the AI data industry is evolving and its significance in AI infrastructure.

“So has this industry moved closer now towards the center of AI infrastructure?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Anna Wyndham:It's clear that it's not just a data industry. It's an industry that transforms human knowledge and judgment and expertise into the signals that AI models can learn from and need to learn from.

0:15Florian:Hey everyone and welcome to a very special episode of SlaterPod today. So today we are not talking about localization, translation, interpreting or anything in that area. Today, we are talking about the data that powers AI. We're talking about the companies that are buying it, the companies that are selling it, and what exactly this entire ecosystem is doing. So who better to have on today than our very own head of research, Anna Windham. Hi, Anna.

0:41Anna Wyndham:Hi, Florian. Hi, everyone.

0:43Florian:So you were in the lead for our new data for AI market report. So let me just give you a quick run through kind of the high level bullet points before I defer to you on the various aspects that we covered in this report. So we, in this report, we sized this market at approximately$9.3 billion. We looked at the data used to train, adapt, align, evaluate, and deploy AI systems. We mapped the buyer landscape, supplier ecosystem, the workflows, workforce, and the market dynamics. And finally, we also found that this industry really has evolved, of course, beyond traditional data labeling, you know, like the type of stuff that companies like Appen used to do, and into a much broader kind of ecosystem focused on making AI deployable in, you know, real business workflows.

1:39Florian:So before we start, I do need to give a quick plug for SlateCon San Francisco, because among the many amazing speakers, we also feature Emily Hsu, who's the head of enterprise AI at ScaleAI, arguably one of the leading providers in this data for AI ecosystem. them. So let's jump right in. And maybe Anna, before we start with kind of the core content of the report, tell us a bit more about your personal background and how this is relevant for this particular piece of research.

2:10Anna Wyndham:Yeah, so I started my career actually in the data for AI market working at Appen, mainly focusing on speech recognition, so speech to text. And that is really where the data annotation or data labeling market began. It's changed a lot since then. And as you said, we've tried to capture in the report what this market looks like now in mid-2026. It's a pretty broad and deep report, but the goal was really to understand what this market looks like now and the extent to which it's changed.

2:48Florian:Got it. So what made us, since later here decide to dedicate a report to data for AI?

2:55Anna Wyndham:Yeah, so I guess many of our listeners are language solutions integrators, and this is an important adjacent market for language solutions integrators. We did an analysis of how many LSIs actually operate in this market, and it's around 10%. But notably, a lot of the players at the top of the market, so the very large super agencies, RWS TransPerfect operate in this and they have significant data for AI businesses and notably fast-growing businesses. And there are obviously synergies between language solutions and data for AI, multilingual expertise and global operations. But as we dug into this, it became really clear that this market has much broader relevance.

3:38Anna Wyndham:So beyond being an adjacent market to language solutions integrators. So this is the market, as you said, that produces the data that is used to build, adapt, align, and evaluate AI models. So in that sense, it's a critical layer of AI infrastructure, and it underpins language AI, so all of the language solutions market, but also AI deployment across the broader economy.

4:04Florian:So we covered this, I remember, I think it was back in 2019, right? When Appen was like, I don't know, a$5 billion company. That's no longer the case because while they had to scale back on some fronts, others kind of exploded. But yeah, we did a report back in 2019. What surprised you the most when you came back to it now?

4:26Anna Wyndham:I guess what surprised me is how much the center of gravity has shifted and how little I recognized, how little was familiar in terms of the basic questions, who buys data for AI? What kinds of data matter most? Where is value being created? What kind of suppliers are in the market and how are they positioning themselves? All of these very basic questions had very different answers when I looked at them this year when we produced the report in 2019. Back in 2019, the main question was scale. So all of the focus was around these data annotation workforces producing data that was used to develop specialized task-specific models.

5:12Anna Wyndham:So we're all familiar with machine translation models, speech-to-text, and there were other areas that were generating a lot of demand, robotics, image recognition, and so on. So going into this research, I kind of expected to find an evolved version of that market. and because of the, I guess, the LLM, I don't know what we want to call it, the LLM explosion or the LLM moment in 2022, I was also expecting to see that there would be a whole lot of new demand for the unstructured large-scale data that we know is used to underpin large language models. But instead, what I found is that the conversation has moved much further downstream, so much further into the question of practical AI implementation into the question of, yes, we have these large language models, they have great broad general purpose capabilities, but how do we actually make them reliable, safe, controllable, usable in real world environments?

6:14Anna Wyndham:And that question changes almost everything about what is needed in terms of data. It changes the type of data being purchased. So we need data that will align, adapt and evaluate models. It changes who's doing the work. So instead of relying on those large pools of generalist annotators, there's now a new and increasing need for domain experts who can provide judgment that will shape model behavior. And we see, you mentioned Appen, we see leaders in this space repositioning around this need as well as new companies emerging that focus primarily on this. It's also changed who's buying. So before the demand was largely organized around applied AI, so machine translation, computer vision, defense applications.

7:04Anna Wyndham:Now, of course, we have demand coming from the frontier labs, so open AI, anthropic, coherent, so on. We also have demand coming from all of the enterprises across every industry who are deploying AI and aiming to extract value from AI. and as well we have demand coming from all the companies that are building their products and trying to make their products use case specific and trying to make their products differentiated. One implication of all of those changes is that the economics have changed significantly so in the past this kind of data work had a very low value per data point but today some of the most valuable data comes from experts whose knowledge has been built up over decades.

7:51Anna Wyndham:And that means that in those cases, the value per data point is dramatically higher. So in short, old world, we had a market focused on building model capability. New world is a market focused on deploying AI.

8:06Florian:It used to be like somebody annotating like bird, car, bicycle-ish, right? And now it's like, hey, does this physics equation look okay?

8:16Anna Wyndham:Exactly.

8:16Florian:So are there still like misconceptions out there about the data for AI market generally, maybe among some of the companies that have heard about it, you know, on our podcast or in other areas and, you know, kind of covered it superficially or there's some kind of broader misconceptions around there?

8:35Anna Wyndham:One of the big misconceptions or misperceptions is that the market is still just that kind of data labeling that you just mentioned, looking at an image, labeling it cat or looking at a social media post and labeling it as having a particular sentiment or so on. another aspect that gets a lot of attention in the media is that is the the need for this large-scale data to pre-train large language models and so there's also a perception that this industry simply provides that kind of unstructured web scraped data and to be clear both of those types of data demands still exist and are still important parts of this market but what's interesting and what's new is this newer layer of demand that doesn't get as much attention but is really the most strategically important type of data now which is the data that makes it possible to release and deploy models so as you mentioned an example would be say a medical physician looking at output to assess whether the response is not just medically accurate, but an appropriate response to a particular prompt.

9:52Florian:Let's talk about the market size. Obviously, people are very interested in that. We size it at 9.3 billion, which given the numbers that we're hearing now with like Anthropic and OpenAI, it seems kind of cute. You know, you have like these like trillion dollar numbers being thrown around for, you know, valuations when these companies go IPO. So what does the 9.3 billion include? What does it exclude? What do we capture with that?

10:16Anna Wyndham:We're capturing the external commercial spending on data for AI. So the Frontier Labs actually have a large component of internal data AI work that's being generated internally and managed internally without the need for external providers. So this is not included in the market estimate. what we're counting here is external spending from frontier labs, enterprises, AI product builders on data sets. So all of the spectrum of data used to build, to adapt AI models, to align AI models with specific behaviors and policies, and all of the expert work that is required to test the models to see where they fail and to evaluate the models to see if they actually do what they're meant to be doing.

11:02Anna Wyndham:So this ranges from, say, off-the-shelf data sets, customer-created data sets, managed data services and the provision of those experts, the AI data platform, so the technology layer. And it also includes licensed data sets. So there's an increasing need for data sets coming from, say, publishers or media archives that can be used either to pre-train large language models or to adapt them to specific domains.

11:37Florian:I guess one there would come to mind is we had Casper Grothwell on the podcast from Oxford. so yeah they would they will probably be in the licensing part of the market yeah and perfect

11:51Anna Wyndham:example of a provider that's really ridden this trajectory of change very fast so they they were previously providing those off-the-shelf data sets and now they're involved in discussions around do we want to license this material for repurpose this material as ai training data so So as well as not including what's being done internally at the Frontier Labs, there's obviously a lot of discussion around how enterprises use proprietary data to make their models deployable and useful and relevant. So this is a very important type of data as well, but it's internal. It's not a case of external spend, so it's not included in our estimate.

12:32Anna Wyndham:And even though that figure may look maybe dwarfed by some of the figures being thrown around by the Frontier Labs, it's much bigger than many traditional annotation market estimates. And that's because we're looking at something much broader than just traditional annotation.

12:48Florian:So take us into the production here of something like this 160-page piece of research. Like, how do you go about something like this? You know, who do you talk to? What do you analyze?

12:58Anna Wyndham:Well, in terms of the scope, the goal was to map the full market to understand the data categories, the demand associated with each data category, all of the different types of suppliers, buyers, and then how value is being created. So like the economics associated with different data types. So we interviewed stakeholders across the system. We spoke to Cohere. We spoke to many of the major data for AI providers. So scale, Appen, LXT, Telus. And we spoke to a lot of the language solutions integrators that have built AI data businesses. So RWS, TransPerfect and so on. So these conversations were so, so interesting.

13:39Anna Wyndham:I can't tell you how many times I've gone back to those conversations. So thank you to everyone who spoke to us. But it's really just such a dynamic and interesting space at the moment. And the other thing that I looked at in detail were the system cards and safety reports and model cards from the Frontier Labs, really super rich source of information. And they really validated what we were hearing from the providers, which is we need these human experts to take part in these alignment adversarial testing and evaluation processes. And that this is a really core need to building models and making sure that they're actually safe, they're actually useful.

14:23So, yeah.

14:24Florian:So one of the biggest themes in the report was the rise of something that we call deployment data. So what is deployment data and why does it matter?

14:32Anna Wyndham:We're using deployment data as an umbrella term to refer to everything that adapts a model after pre-training. So for pre-training a large language model, you need the massive unstructured data. And for under-deployment data, we're including the data used to adapt models to specific domains. So, for example, to make them specific to legal or medical or financial contexts, we're including the data that's needed to align their behavior. So this is data that gives models examples of how to give a good response or what does an enterprise actually need when someone asks, can you write me a marketing plan or analyze this Excel data?

15:25Anna Wyndham:and it also includes the adversarial testing. So this is where experts will go in and try to expose the failure modes and the weaknesses in the model and try and get it to go against policy or provide dangerous information or to provide inaccurate information. And finally, deployment data includes all of the data, so the gold data sets needed to evaluate the model. So in order to understand if the model's working, first you need to define what good looks like. To define what good looks like, you need experts. You need journalists. You need medical experts, financial experts to say this is what good looks like.

16:08Anna Wyndham:And then you need those experts again to evaluate the output of the model and say, yes, in this case, it aligns with what good looks like or it doesn't. And an interesting feature of all of these types of data or these processes is that they are very iterative. So the more you test and evaluate the model, the more you expose or reveal or unearth the ways you need to change it. And in order to change it and adapt it further, you need more data. So it's very much a cyclical process.

16:40Florian:Let's talk about, you mentioned that red teaming and also the alignment. So what does that work actually look like? What are people actually doing there?

16:48Anna Wyndham:Yeah, so I'll start with alignment. So this is how data is used to teach models how to behave. So it shapes things like healthiness, safety, the tone that is used, policy compliance, and when a model should answer versus when it should refuse. So alignment teaches a model how to act. So when you're interacting with a model and it refuses to answer a question, this is as a result of this alignment process. So this covers a wide range of tasks. You might have experts who write ideal responses. They might be ranking outputs to say this output is better than this output. the experts might be looking at whether or not outputs align with policy or providing examples of how an output should align with policy.

17:41Anna Wyndham:So it's about teaching models to follow specific instructions or formats. So imagine a prompt who asks, how do I treat chest pain? And you get two possible answers. One suggests a diagnosis. One is much more nuanced. explaining possible causes, highlighting uncertainty, recommending the user seek medical advice. So the human contributor working on this data task would be tasked with selecting the safer and most appropriate response. And the contributor's decision there becomes a training signal for the model. So this is through various training techniques used to update the model. So you can see why expertise is important here, because a generalist simply wouldn't be able to make a good judgment in that case.

18:32Anna Wyndham:The other one you mentioned is adversarial data or red teaming.

18:36Florian:Red teaming, yeah. Did you know the origin of that term, red teaming? Is that like a military term or something?

18:42Anna Wyndham:Yeah, I think it is a military term. So it refers to when, I guess, forces pretend to act like the enemy and try to find weaknesses or try to infiltrate. I think that this is where it comes from. But it's a really interesting term and it's all the way through the model cards and safety cards because this is how you find the weaknesses and failure modes. So red teamers are designing prompts. They're trying to bypass the safeguards. They're trying to exploit flaws in reasoning. They're trying to trigger policy violations. and they're trying to expose inconsistent behaviour across languages, domains and use cases.

19:26Anna Wyndham:And you can see again why it's so important to be a domain expert in these cases because as a generalist you may have some ideas about how you might be able to get around the model's safeguards but you really need to understand a domain to understand the safeguards firstly and then to detect when they've been bypassed. So just as an example of the breadth of human expertise involved with red teaming, OpenAI reported that the team for GPT-4, the red team included 100 external red teamers. They spanned domains from cybersecurity, healthcare, law, finance, psychology, misinformation, child safety, linguistics, across 45 languages across 29 countries.

Read the full transcript

20:13Anna Wyndham:So that's an example of the type of demand and you can also see why multilingual, where multilingual expertise comes into it as well.

20:21Florian:Yeah, because you could arguably maybe exploit, like, I don't know, the prompt is in a different language or something and that could maybe be one attack vector. So the report places a lot of emphasis also on evaluation. So why is evaluation becoming so important or is already so important?

20:42Anna Wyndham:This was also a big theme at SatoCon Evaluation, where we had a presentation from RWS and co. here, and they were talking about how they actually evaluated the model that they built together and how important it was to have a custom evaluation benchmark because you need to define what good looks like in your specific case before you can evaluate against that. And evaluation is probably the most strategically important category of all. Without evaluating if models do what they're meant to do, it's not possible to release a model. We've seen a few frontier labs holding back models for various reasons, in some cases possibly because the evaluation results were not what they needed.

21:32Anna Wyndham:The other thing evaluation gives you is that it's the evidence that you need to provide to the downstream users of AI to say, yeah, we tested it, we got these results, and it showed this, so it does this. Without that transparency, it's not possible to deploy AI. So this is kind of like the control system for AI development. and you can see that this is where the relationship between the frontier labs and the domain experts becomes important because the frontier labs are building the model but the experts are telling them the standard against which the model should be judged. So an example of a data task might be a physician who creates realistic clinical scenarios to test whether a model provides safe medical guidance or a legal expert assessing whether a model correctly interprets regulatory obligations.

22:25Anna Wyndham:So as well as evaluating the model, the experts are involved in defining the, I guess, the terms of the evaluation as well. And what's interesting is that the frontier labs are very protective of their evaluation data sets. So they're created with these experts. They're very, very meaningful because they define what is being tested. They're defining what's important and they're and they provide the evidence that they enable the extraction of evidence that the model is performing in certain ways so this is where we see this kind of paradox which is that the more capable the models become the more important it is that we have humans who can tell us in what ways they're actually useful, if they are trustworthy, if they are safe.

23:21Florian:So let's talk about those humans. So what's changing about the workforce behind AI? I mean, we spoke about the whole shift from annotation to much more involved and expert-driven kind of activities. But yeah, tell us a bit more about the workforce here.

23:35Anna Wyndham:The workforce kind of relates to those different layers of data demand that we already talked about. So there is still demand and need for these large labour pools of generalist annotators. So in that case, the kind of challenge, well, the focus up until a couple of years ago was scaling those resources. So sourcing them, scaling them, ensuring that people were well trained, ensuring that the output of their annotation tasks was quality checked. And that's still important. But because of this new, I guess, bottleneck to deployment, which is ensuring that a model is safe and having evidence that it is safe, that creates this demand for human experts.

24:26Anna Wyndham:Increasingly, the most valuable contributors are domain experts. So we see Frontier Labs reaching out to external data for AI providers, companies like Appen and Scale and so on, and saying we need physicians, lawyers, engineers. One example we found was Frontier Labs recruiting investment bankers for hundreds of dollars an hour to create evaluation frameworks and rewrite model outputs to reflect how senior bankers actually think. so yeah there are these different layers of demand still the generalist annotators are very much in demand but on top of that specialists and experts.

25:09Florian:So what does it take to work for these frontier labs right what would separate a winning vendor from others that don't get through because these companies have infinite money at the moment so how do you work for them?

25:22Anna Wyndham:Well this is obviously where the major suppliers are competing. And Frontier Labs, according to the interviews that we had with various suppliers, Frontier Labs really need providers that can solve novel problems because this process of AI development is still one of discovery and finding out. And their data needs one week might look completely different next week depending on shifting goals or the results of adversarial testing and eval. So one week they might need to build multilingual red teamers across different languages. Another week they might need security experts to probe model vulnerabilities.

26:07Anna Wyndham:They might need companies that can support with simulating different data scenarios. So, for example, simulating a doctor-patient conversation, And this can be for various reasons to ensure that for data privacy reasons, for example. So you need examples of those conversations, but you can't use the originals. So you might simulate those conversations. And our conversation with Data Mundi, who is a language solutions integrator who pivoted into this market, said that they found themselves folding napkins 100 times a day and recording that process in order to teach models how to fold napkins. And RWS also referred to doing quite what they called wacky tasks.

26:55Anna Wyndham:So really interesting tasks, but I guess it's the flexibility that the Frontier Labs need. Another thing that stood out is the pace. So Frontier Labs are constantly experimenting. And we spoke to Steve Nemza from Telus Digital, and he said that they'll need a provider for a fast, short-term project without much lead-up while they're testing ideas. As soon as they discover something works, suddenly they need to scale really fast. So the providers that are able to first respond to those quick short-term needs and then move into scaling other ones who are winning. But probably the biggest differentiator is trust.

27:37Anna Wyndham:So increasingly, there's this focus on is data governed? Is it traceable? Can we defend the whole process that it came through to get to us? So they need to know where the data came from, who produced it, which contributors were involved and how they were vetted the rights around the content. They need to see that the process withstands scrutiny.

27:59Florian:Other than the major frontier labs, I guess there's like two at this point, right? OpenAI and Anthropic. And then you have, you know, Mistral and Cohere, et cetera. But who are the major buyers today? And like, what are they actually buying? We've touched on some of this, but yeah.

28:15Anna Wyndham:Yeah, as you said, there's only a very small group of frontier labs. And you have your top tier and a couple of others just behind in terms of size. So though they are few in number, they are spending very heavily and increasing their spend on data. So that is where the demand is concentrated. But apart from that, we have enterprises and AI product builders. And across enterprises, certain industries are much more focused on AI operationalisation, AI deployment, and have greater data needs. We go into that in the report, breaking down which industries are more forward and which ones are lagging in terms of AI adoption.

29:00Anna Wyndham:And the other notable biotype, which we count as a different biosegment, is the rise of sovereign AI programs. So these are national AI initiatives. They want to have greater control over all AI infrastructure. They don't want to be using infrastructure from the US. So this is creating a new wave of demand. And all of the demand that we've previously seen coming from the US is kind of being myriad or replicated, but across different languages now. So they need, obviously, language coverage and they have specific security requirements. Something interesting is that all these different buyers need very different things and they evaluate suppliers very differently.

29:48Anna Wyndham:So we spent quite a lot of time unpacking those differences in the report. So you can check the report for procurement patterns and buying behavior, how they select their suppliers and so on in the report.

30:01Florian:There's some amazing tables in there that if I was a salesperson in this industry, I would want to get my hands on. So, you know, do go there. So for the LSIs, is this like a genuine adjacent market or a fundamentally different business? We had Paul Carr from WeLo Global, formerly WeLocalize, on the podcast. And I think he was a little bit more on the, this is quite different, not fundamentally different, but quite different. So what's your take?

30:28Anna Wyndham:We interviewed a wide number of LSI's. And so they were very clear that the capabilities and resources that LSI's have from delivering language solutions are a major advantage when moving into this market. But they also, as Paul Carl mentioned, a significant degree of transformation is also needed to succeed. So on the one hand, the assets or the strengths that LSI's have are the multilingual contributor network, experience managing distributed workforces, of course, expertise in language evaluation. A lot of these AI building processes are moving deeper and deeper into the everyday of the language solutions integrator segment.

31:15Anna Wyndham:Many LSIs told us that they see it as an extension of what LSIs already do, which I thought was an interesting, really interesting way of seeing it. We spoke to Thomas Burkett from RWS Train AI, and he said, it's really an extension of what LSIs have already done. It's about enabling communication across languages and cultures and helping AI systems to understand language and operate globally. It's just an extension of the same mission. So I thought that was interesting to see it as an extension of the same mission. And you can see where the advantage comes in. If a buyer needs German physicians, Japanese lawyers, Arabic financial experts, LSIs have the capabilities for that.

31:56Anna Wyndham:But as you mentioned, there are lots of differences. It's a different operating model. There's different expertise needed internally. And one theme that came up is that it's really critical to have people internally that can speak the language of machine learning, people who can engage credibly with frontier lab researchers, understand the objectives, translate those objectives into operations, because that relationship is very collaborative and iterative. It evolves as it goes along. So the frontie labs really need to have people on the supplier side who can understand what they're trying to do and respond to that.

32:38Anna Wyndham:And I think as Paul mentioned, the tooling is very different. So the technology and infrastructure you need are very different. So it is a genuine transformation. LSI's have some advantages, but they also need to develop new capabilities.

32:51Florian:So let's close on a question I have. Let's call it an industry. Let's say the data for AI industry. So has this industry moved closer now towards the center of AI infrastructure? And what role does this industry play in kind of the broader AI economy, which is blossoming and exploding at this point?

33:12Anna Wyndham:Yeah, well, I mean, data, this industry has always been a support industry and it used to be referred to as the picks and shovels industry, which is a bit kind of an unglamorous way of seeing it. But I would say that the industry has moved much closer to the centre of AI infrastructure because these data tasks which are aimed at making models safe and deployable in real world scenarios and suitable for specific contexts and use cases, this is really critical if AI is going to be deployed widely and if it's actually going to be useful, if companies can actually extract value from it. So in that sense, the industry is helping solve one of the biggest challenges in AI, moving to real world deployment.

34:04And I guess the other big takeaway is that it's clear that it's not just a data industry.

34:12Anna Wyndham:It's an industry that transforms human knowledge and judgment and expertise into the signals that AI models can learn from and need to learn from. So that's why it's moved much closer to the center of AI infrastructure. It's the difference between building a capable model and deploying a trustworthy and usable AI model. Great.

34:34Florian:Well, that was a sneak peek of 30 minutes into the Data4Rile report that we have on our website for subscribers. You can download it for those who want to buy it as a one-off, head over to the website and get it. I think it's a phenomenal piece of research that everybody in the broader language industry, language solutions and language AI market should go and read because it does represent a massive growth opportunity if you are serious about it. And I think this report is definitely a great starting point to learn about this market. So, Anna, thank you so much for taking the time today.

From the publisher

Slator's Anna Wyndham joins Florian on the pod to discuss key highlights from the Slator Data-for-AI Market Report, which sizes the global market at USD 9.3bn and examines the ecosystem supplying the data needed to train, adapt, align, evaluate, and deploy AI systems.

Anna explains how the market has evolved far beyond traditional data labeling. While annotation and large-scale training data remain important, she argues that the market’s focus has shifted toward helping organizations deploy AI safely and effectively in real-world settings.

Anna highlights the growing importance of “deployment data”, data used to adapt models for specific domains, align behavior with policies, conduct adversarial testing, and evaluate performance. She notes that these activities increasingly rely on subject-matter experts, creating demand for professionals such as physicians, lawyers, engineers, and financial specialists.

The discussion also explores how frontier AI labs, enterprises, and sovereign AI initiatives are driving demand. Anna shares that buyers increasingly need trusted providers capable of sourcing expert talent, scaling rapidly, and maintaining rigorous governance around data provenance and quality.

For language solutions integrators (LSIs), Anna sees both opportunity and challenge. Existing strengths in multilingual operations and workforce management provide a natural advantage, but success requires new capabilities, including expertise in machine learning workflows and AI evaluation.

More from SlatorPod

All 39 episodes
#287 How the Market for AI Data Has Become a Major Growth OpportunitySlatorPod · 35 min
Listen in VO