#298 Ryan Kolln: How Appen Trains the World's Most Powerful AI Models

6 Nov 2025 · 51 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: Eye On A.I. - Episode #298

Episode Overview

  • Host: Craig S. Smith
  • Guest: Ryan Kolln, CEO of Appen
  • Episode Title: How Appen Trains the World's Most Powerful AI Models
  • Release Date: [Insert Release Date]
  • Sponsor: AGNTCY - [Website](https://agntcy.org/)

The episode focuses on how Appen trains and evaluates AI models through a rigorous evaluation system that incorporates human feedback, cultural nuances, and advanced methodologies.

Key Topics Discussed

  1. The Role of Evaluation in AI Development
  2. Importance of Evaluation:
  3. Evaluation helps measure the effectiveness of AI models beyond static benchmarks.
  4. High benchmark scores do not guarantee good user experiences.
  • Human vs. AI Evaluation:
  • Human evaluators play a crucial role in evaluating AI models.
  • Evaluation systems must adapt to cultural contexts to ensure relevance in global markets.
  1. Cultural Nuance in AI Models
  2. Localization:
  3. Different cultural contexts significantly affect how AI models perform and are perceived.
  4. Example: Parenting advice varies widely across cultures (e.g., US vs. Middle East).
  • Sovereign Models:
  • Growing interest in sovereign AI models that cater specifically to local cultures.
  • Appen collaborates with model developers globally to address these needs.
  1. The Evaluation Process
  2. Human Annotation:
  3. Human feedback is critical for improving model performance and retraining.
  4. Evaluators assess factors such as accuracy, format, bias, and cultural appropriateness.
  • Feedback Mechanisms:
  • Evaluators provide both preference judgments and qualitative feedback to improve AI responses.
  1. Technological Infrastructure
  2. Evaluation Stack:
  3. Appen has developed a robust technology stack for managing human evaluators and their tasks.
  4. Real-time quality checks using AI complement human evaluations.
  • Crowd Management:
  • Appen’s database includes millions of potential evaluators from diverse backgrounds.
  • Efforts are made to ensure evaluators are qualified and capable of delivering high-quality feedback.
  1. Market Dynamics and Future Trends
  2. B2C vs. B2B Markets:
  3. The B2C market is dominated by a few large players (e.g., Google, OpenAI), while the B2B market is fragmented with niche applications.
  • AI in Evaluation:
  • AI models are increasingly being used to evaluate other AI models, though human evaluation remains the gold standard due to the complexities and subtleties involved.
  • Future Challenges:
  • Significant challenges remain, including the need for continuous model improvement and the balancing act between local relevance and global applicability.
  1. Economic Aspects of Human Evaluation
  2. Compensation Models:
  3. Appen pays above minimum wage and adjusts rates based on the complexity and specialty of the task.
  4. Opportunities exist for individuals to earn a significant income depending on their expertise.
  1. Continuous Learning and Model Improvement
  2. Post-Training Evaluation:
  3. Emphasis on evaluating AI post-training to ensure that models adapt and respond effectively to human interactions.
  • Future Innovations:
  • Potential breakthroughs in architecture and methodologies may lead to significant advancements in AI capabilities.

Key Takeaways

  • Human evaluation is indispensable for the development of effective AI models, particularly in capturing cultural nuances and providing nuanced feedback.
  • Appen’s structured and diverse workforce allows for comprehensive model evaluation across various languages and cultures.
  • The AI evaluation market is evolving with increasing integration of AI tools for assessment, but human input will remain critical for high-quality outputs.

Conclusion This episode of Eye On A.I. delves into the complexities of training and evaluating AI models, highlighting the importance of human evaluation combined with technological support. As the AI landscape continues to evolve, understanding cultural nuances and maintaining quality control through both human and AI evaluators will be key to success.

Further Listening

  • Follow Craig Smith on X: [@craigss](https://x.com/craigss)
  • Follow Eye on A.I. on X: [@EyeOn_AI](https://x.com/EyeOn_AI)

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The novel in the LLN space is you've got a large number of models, all serving a similar purpose, which is to bring a lot of context and knowledge and knowledge and knowledge. of the information we have into a way that can be easily summarized and presented to humans. So I think for the first time, there's this real reliance on benchmarking to understand which models perform any better than others. But I think as we've seen through some of the benchmarking, even though there is a particular high score on a benchmark, that doesn't always translate into a good user experience. For most of our customers who are operating at a global scale, the evaluation plays a really important part into bringing that cultural nuance and context of how the world operates.

0:48Build the future of multi-agent software with Agency. That's A-G-N-T-C-Y. Now an open source Linux foundation project, Agency is building the internet of agents, a collaborative layer where AI agents can discover, connect, and work across any framework. All the pieces engineers need to deploy multi-agent systems now belong to everyone who builds on agency, including robust identity and access management that ensures every agent is authenticated and trusted before interacting. Agency also provides open, standardized tools for agent discovery, seamless protocols for agent-to-agent communication, and modular components for scalable workflows.

1:45collaborate with developers from Cisco, Dell Technologies, Google Cloud, Oracle, Red Hat, and more than 75 other supporting companies to build next generation AI infrastructure together. Agency is dropping code, specs, and services, no strings attached. Visit agency.org to contribute. That's A-G-N-T-C-Y dot O-R-G. Hello, my name's Ryan Collin. I'm the CEO at Appen. I've been at Appen for over seven years now, focusing on growing the business in the exciting world of AI data annotation. Prior to Appen, I have a background in engineering, specifically in data requirements for telco networks. They've done very large-scale data annotation and modeling around network optimization for telco networks.

2:51But very excited to be here today and talking about how Appen's supporting the development and evaluation of LLM networks. Can you start for listeners that aren't that familiar with the space? I mean, for a long time, LLMs, well, I don't know this to be true, but it seemed like everyone was relying on performing to certain benchmarks. But quickly, people realized that that's a very narrow way to evaluate LLMs. And it can certainly be gamed. you can train to meet those benchmarks or exceed those benchmarks. Can you talk about kind of the evolution of evaluation? And as I said, I'm particularly interested in the role of human evaluators.

3:49Yeah, it's a very interesting market. And I think, you know, what's become, you know, novel in the LLN space is, you know, you've got a large number of models all serving a similar purpose, which is to bring a lot of context and knowledge around all of the information we have into a way that can be easily summarized and presented to humans. So I think for the first time, there's this real reliance on benchmarking to understand which models perform better than others. But I think as we've seen through some of the benchmarking, even though there is a particular high score on a benchmark, that doesn't always translate into a good user experience.

4:37And I think the reality is, is the benchmarks, they're static and typically quite narrow. Therefore, it's really difficult to get one benchmark that, even within a narrow domain, that covers all of the edge cases. I think the other issue is that the ways that a human use these models evolve very rapidly, the model itself evolving rapidly. so while the benchmarks play a really interesting snapshot into model performance i don't always think that it's the best way to really get a measure of how good the user experiences for the models that leads us into both llm evaluation of other llms in human evaluation can you start with human i mean that that was part of the original uh development of of these llms with reinforcement learning with human feedback uh but it's gone beyond that hasn't it yeah that's right and um even before the llm uh model developments over the last kind of five to eight years or so So evals for large, human-based evals for large models has been going on for a very long time.

5:56And Appen has played a big role in this. So there was a type of human annotation called relevance. And really what it was is providing feedback, subjective feedback into how the large models run by the very big search and social media companies were performing. And these are the core models that are related to for a given search term, what are the right search results that are being shown or for a given user experience, what is the right ad that is to be shown. And human feedback into the model development lifecycle has been operating at really large scale for a long time, but it's a small niche within a few very large technology companies that's been riding on that.

6:46And it's probably helpful for me to explain what an evaluation looks like or what search relevance looks like. And it's a fairly simple but subjective task, which makes it complex. So at a really high level, for a given prompt, here are two responses. Which one do you prefer and why? right and the type of questions that we ask our annotators and our workforce really are dependent on what the customer is trying to do sometimes they'll be looking for how accurate was the response was it format in a right way that is you know can be easily digested is it grammatically correct and then you get into you know was it harmful was there any biases contained so you can think about all of the different subjective elements that a model builder would want to know how well their model's performing is up for grabs in these evaluation tasks.

7:48And then what can be really important also is typically there's a follow-on question, why did you think that way? Or what would have been a better way to answer that question? And it serves two purposes. One is that it's used to measure the performance of how the models are performing, but also it's used oftentimes in the retraining process to improve the model performance going forward. Now, what is the unique part here is for most of our customers who are operating at a global scale, the evaluation plays a really important part into bringing that cultural nuance and context of how the world operates.

8:35So I'll give an example which helps highlight this. Recently we're asked to find people in our network who were experts in parenting advice, right? And if you find someone who's an excellent, you know, parent in the US, they may give very different parenting advice to what is done in the Middle East or in India or in Africa. So there is this cultural context into the evals that is incredibly important. So the models work and are accurate and are giving great advice, not just based on a Western point of view, but on a real global point of view, and it brings in that cultural nuance. That's interesting.

9:18And does that then train the model to recognize what culture the customer, the person interacting with the model is coming from so that it answers in a culturally relevant way? Or is it averaged out, which seems to me may be problematic as well because you end up with just extremely general, not very useful responses? Yeah, the cultural context is really important. And the idea is that it is addressing the requirements within a culture. But then there is also a base overlay around what is the right thing to be doing at a humanity level. So it is trying to get the balance right between how do we want the models to be behaving at a generic view across all cultures, but then capturing that nuance within a specific geographic or demographic setting.

10:21Yeah. Are you working with sovereign models in countries outside the West so that they're trained to their cultural viewpoint? I mean, that's, of course, one of the big complaints among smaller countries that they're getting models that are trained to a U.S. viewpoint. Yeah, it's a really good example. And the rise of the solid models are really addressing this issue that I've spoken about, where it's like, it's the cultural context isn't there. There's an interesting set of developments. And, you know, we work with some of the specific model developers in different regions. I think for our side, though, where the bulk of our customers are the large technology companies.

11:14But we do have Appen is unique in the industry where we work with both the very large US LLM model builders and the Chinese model builders, which represent kind of the most innovative model builders on the planet. And there is and we do see some differences in there. And it's not always cultural. I mean, that's a big part, but we also see it's derived by the distribution models. So in China, we see more work related to medical and financial advice, and that's driven by the super apps where you go in and it will give you medical advice to a consumer. and that's a really big part of the ecosystem there.

12:06Where in the West, you don't typically go to your smartphone for medical advice. At this stage, you go to your GP. So there is a different also set of data requirements and set of evals based off how consumers use technology in different environments also. So does that mean that with the human evaluators that you have teams from each culture that you're working in, you're not using a global team or a US-based team to evaluate or tune Chinese models? Or are you? I don't know. Yeah, so there are two parts to what our customers in China want. So there is support for the domestic applications, and then a lot of our customers in China, technology companies, have global customers, right?

13:01And, you know, there's some really good examples on social media, also on e-commerce, that there's a very big global footprint. So we're helping both with the domestic and the international data requirements for the China customers. You know, OpenAI, and it's the one that comes to mind, But certainly in some of the other models, consumer-facing models, periodically you'll get two answers and be asked to choose one or the other. To me, that's a pretty flawed way of, because a lot of times I'm busy and I just pick one just randomly because I don't want to have to read both. uh yeah is that useful that consumer facing uh uh effort to to have them show a preference okay i think it's um you know when looking at the evaluation of the models and how they're working and the data that's used to retrain it's um there are many many different sources that are used All right.

14:16The benefit of going with an approach like what Appen provides, we are qualifying people. We know their exact demographic backgrounds, their exact interests. We know that they're going to sit down and pay attention to that task and not just click option A every time. So there's a lot of precision in the data that we're providing, albeit with very subjective feedback. And I think that, you know, looking at, you know, the model builders look at things like the example that you gave about OpenEye, but also engagement rates, click rates, the way that they respond to the question. So there's an ensemble of data inputs that are provided to improve the performance.

15:01And what companies like Appen provide is a really important part of that. And what you're providing is a very focused, initially, I mean, we'll talk about evaluation of other elements. But on the human side, you're providing, as you said, a focused and trained workforce with instructions of what they need to be evaluating. How many people do you have in your database of potential evaluators? It must be massive, or maybe not. I mean, maybe that's my other question. How many people do you put on a model? Yeah. So we have millions of people in our database and getting access to people isn't typically our challenge.

15:53It comes through finding the right people who are going to deliver the demographic requirements for what is required by the customer and then the ability to ensure that they are creating data that meets the requirements of the customer. And I'll give an example. If we had millions of people and they were only based in the US, we would have a very US-centric based workforce. The benefit that Appens brings is we've got people in, I was just looking the other day, I think it was one country where we didn't have someone in that database on the planet. and we've been doing this for a very long time and and the the art of what we do is running a project where the customers were asking us to you know connect with people in a hundred countries in parallel let's speak a hundred different languages and have a hundred different um you know ways of interacting with us and payments and support etc but it's that the customer wants a common framework around how they're going to, the type of questions they want the contributor to be asking, how they're going to understand whether their output is high quality or not.

17:09So the ability to translate this single task into 100 different cultures and 100 different languages and then ensure that the workers are really attuned to the nuance of the task that's being asked of them. It's a non-trivial task and, you know, something that we do and we productize and we have a great set of infrastructure to be able to do that. Yeah, and that's interesting. So you have human evaluators evaluating the evaluators. That's right. So, yeah, and quality is a really interesting topic for human data. All right. So when we have someone evaluating the model of performance, we need to check that they're not just choosing option A every time.

18:04All right. And there are many different techniques that we've got to ensure that they're providing the right kind of effort and the right quality into that data. the other issue we face a lot Craig is that um you know there are people who um try to be in a different country to what they are or try and you know say that I've got these skills when they don't so we have a pretty robust process highly automated to verify if people are in the location where they are are they humans um do they have the expertise that they want um and that's that's again, some of the things that we've built up over doing this since, you know, over 20 plus years, we've got a lot of expertise that we've built into our products to make sure that we're getting high quality data.

18:53And typically, to evaluate a customer's model, how large of an evaluation team are you fielding? Yeah, it's a good question. It does vary. And it varies about the task type, but it varies about the global reach. But it's not uncommon for us to have thousands of people on a specific model eval working on an ongoing basis at any given time. So the numbers can get very large very quickly. We also see there can be bursts of data. You know, a customer recently asked us to find 20 ,000 people to evaluate the performance of their model. And that's a lot of humans. Yeah. And so Appen has this database of people that you've recruited and put through some evaluation, both automated and human.

20:06And then when you want to pull together 20 ,000, I imagine a lot of that's automated. But how then do the people interact with the model? Do you just point them at the model or do they interact with the model through an app and website or app? Yeah, great question. So there's two parts to our technology stack. One is you can think about as workforce management. And that's where we have the database of crowd workers, the ability to go and verify and put them through different qualification steps to understand what they're going to be good at. And also to advertise the tasks to them. So we have a crowd gen is our brand from our crowd and people go in there and they see what tasks are available to them.

21:01And if we find a project that matches someone, we will notify them. There's some great earning opportunities for you. So that's one side of it. The other side is the data annotation platform. And this is exactly what you're talking about, Craig, where they come in and complete the tasks. And Append has an exceptionally flexible and powerful data annotation platform. and the reason why it needs to be flexible because every task that we get from our customers is a little bit different. The question might be different, the scoring, et cetera. So we need to create very modularized components for our crowd workers to come in and complete the task.

21:46The other thing what's really powerful about the annotation platform is the ability to check for quality in real time. So I'll give an example. If the task is, you know, evaluate the performance of this model and rewrite the response in numbered bullets, so 1, 2, 3, 4, 5. We have a lot of AI checks that run in parallel to what the contributors are doing. And it could be as simple as, was the formatting correct? Was it on topic? Yeah. Was, you know, we get into specific things like spacing and, you know, grammatical checks. So there's a lot of AI that we're running in parallel to the task to make sure that the quality is where we need it to be.

22:38And Appen is publicly listed in Australia. When was it founded? Yes, it founded in 1996. And the founding story is an interesting one. and we've been doing AI work since the start. So there was a company in 96 called Nuance. I think they got acquired by Microsoft at some point. They were building a speech recognition model where you could call up and place bets on horse races, right? And they built that in the US and this was in 96, so it wasn't called AI back then. and they thought that Australia would be a next great market to roll this out to. They rolled it out in Australia and it didn't work because it couldn't understand our accents.

23:29So they called up our founder, who was a professor of linguistics at the time at a university in Australia, and said, can you go and help us collect Australian-accented English that was used then to train the speech recognition model so people could place bets on horse races. That's fascinating. Did she, it's a woman, right? Yeah, that's right. Julie Von Willow. Did she use Mechanical Turk or what? Did she use an existing crowd working platform? No, I don't think Mechanical Turk existed at this point in time. So this was pre the data annotation industry and really, you know, the tip of the spear of the foundation.

24:13And so Julie went out and found people and brought them into a recording booth and recorded them speaking Australian-accented English for the specific requirements of nuance at the time. And, you know, we did some really great work early on and that snowballed into, you know, really supporting the development of speech recognition models. We did a lot of work with IARPA and DARPA following September 11 around building up models that had a specific focus on Arabic languages and all of the different dialects. You know, it's interesting about Arabic is that depending on the location, it can be spoken very differently, but it's all written the same.

25:03So Appen developed a way and it was very innovative to phonetically transcribe colloquial Arabic that was then used to build into the speech recognition models. Yeah. So following that, we did a lot of work with Amazon Alexa early on to support the development there. and really the focus of AppBan, it was fairly focused on speech models at the time and we built out this global network of people around the world to get language and dialect accuracy. And then that's what, because we had this big network all around the world, some very large technology companies came to us and said, hey, that would be great if we could leverage that to evaluate the performance of our global models.

25:58So that was the evolution of Appen from speech into relevance. And then from then, we've done a lot of work in image annotation, continue to do a lot of work in speech, and then more recently, clearly, for Jennifer Bayer. Build the future of multi-agent software with Agency. That's A-G-N-T-C-Y. Now an open source Linux foundation project, Agency is building the Internet of Agents, a collaborative layer where AI agents can discover, connect, and work across any framework. All the pieces engineers need to deploy multi-agent systems now belong to everyone who builds on agency, including robust identity and access management that ensures every agent is authenticated and trusted before interacting.

26:57Agency also provides open, standardized tools for agent discovery, seamless protocols for agent-to-agent communication, and modular components for scalable workflows. collaborate with developers from Cisco, Dell Technologies, Google Cloud, Oracle, Red Hat, and more than 75 other supporting companies to build next generation AI infrastructure together. Agency is dropping code, specs, and services, no strings attached. Visit agency.org to contribute. That's A-G-N-T-C-Y dot O-R-G. I was going to say that on the mechanical Turk thing, that Fei-Fei Liu built the ImageNet dataset that then was used to validate convolutional neural nets and set off the AI supervised learning AI craze.

28:04That spawned a whole industry of data annotators. I was close for a few years with a company called Labelbox, very bright guys who have a platform. And I think I had Appen on the podcast at some point talking about your data labeling system. but uh is did did during that supervised learning uh phase not that it's over but before transformers and generative ai took over was was that a big part of appin's business yeah it certainly grew to be an important part of the business um we did um you know work in computer vision annotation and still do. So certainly the ImageNet and the Fei Fei Li and the deep learning kind of evolution was really great for image recognition models.

29:09There was also neural nets used in the speech model development, which really spurned a next wave of the data collection and annotation for the speech models also. Let's talk about, you built, this a network of evaluators human evaluators you have a system that that they work through to evaluate uh responses from from ai models this is for uh fine tuning is that right or or is it for the foundational training it's mostly for in an llm context for post training so there are um post-training and eval so you know in the llm model development there is pre-training post-training pre-training is hoovering up you know all of the available copy on on the internet and books etc highly curated um but it doesn't bring in the the human context post-training is making sure it's bringing in how the models that have this great knowledge want to be able to convey that to humans.

30:21And is that part of the RLHF cycle or is it after that? It's part of the RLHF cycle. So a lot of the evals, now there's a shift from reinforcement learning with human feedback, which was effectively providing examples around how the output should look to reinforcement learning with verifiable rewards. So this is more evaluating at a rubric base, so with a scoring system, how well the model is performing, which is then used to create a reward model, which guides the output. So there is an eval component that's used as a really critical part of post-training. And then there is evals, which more just like, how well is it performing, which may or may not be used to train the reward models.

31:17I see. And that, that, uh, the reward model, uh, is you're, you're not talking about fine tuning. You're talking about post-training, right? Before the model is released. And, um, the, um, you know, this is, this is a tremendous amount of labor. I remember when I first found out that there were human, uh, annotators for supervised systems. It was like, oh, this is AI, but in fact, there's this army of humans out there working on it. You've moved into LLM evaluation, right, into using AI to evaluate AI. Can you talk about that? Yeah, sure. So there is a big part of the market that's using AI to evaluate AI, right, which is really important.

32:21There are some challenges that that presents, though, because, you know, you are trusting another model to effectively act as a human. and the ability that the the level of confidence that you have in that the uh eval model really depicts what confidence you have in the outcome of those evals right so um there is no perfect eval model at the moment what we are seeing is that there's eval models being created that are very narrow domain specific right um you know one it's cheaper it's faster but it's you know, fine-tuned to be very specific evolve. I think that as we wind the clock forward, AI is evaluating AI is going to be a big part of the market.

33:14I think that human evaluating AI is still going to remain the gold standard and be, you know, a really important part of the market going forward. Yeah. And does Appen have evaluation LLMs that you've developed yourself? Yeah. So we have the ability to bring AI and human. And what becomes really important is an AI will evaluate the AI model and will come with a confidence index. And where it's low confidence, it can be verified by a human. Also, we use the human to do spot checks around making sure that the model is performing right. So the ability for us to incorporate humans with the AI evals in a single workflow, I think that's something we're betting on and investing in and making sure.

34:09I think that's where the market's going to be headed in the long run. And when you say the market for evaluation, is it a crowded market? It's certainly a large market. It's a large market. It is crowded. And what we typically find is that there's competitors who have their different strengths. So there will be some who are very focused on very advanced science and the hard science of mathematics. There's some who specialize in coding. The thing that Appen brings is that broad contextual knowledge with the ability to go domain-specific deep in different areas. So we can find people all around the globe who represent humanity, but then also translate that into people in Thailand who have a PhD in molecular biology.

35:06Now, that pool gets a lot smaller. We can't find tens of thousands of those folks, but it's the ability to go both broad and deep is where Appen shines. How do you figure out what to pay people? Because if you're doing a PhD in molecular biology, certainly they're going to demand a higher pay than somebody who's just evaluating garden variety question and answers. supply demand economics which which drive but there's also when we find someone who is delivering consistently high quality data and has a um a real interest in the work that we do and a passion for it and wants to come and complete that work those folks are incredibly valued both for us and for our customers so um we do differentiate our pricing even an individual level yeah is this something that somebody can make a living on or is it sort of gig work that that uh you know people drop in and out of yeah it's increasingly we're seeing people um who this is a large part of their their earning opportunity so we have folks that happen who you know 20 plus hours a week and they're consistently doing this work and again are really passionate about it and really enjoy doing the work and understand the value that they're providing to the AI ecosystem.

36:36So I think the human data annotation industry gets a bad rep sometimes and it's pushed to the lowest common denominator. When you speak to our crowd workers, they really enjoy the work that they're doing and they're rewarded appropriately and they're super passionate about the industry. Yeah. Do you, have you, I mean, there, there are certainly, uh, in India and the Philippines, for example, uh, a whole industry of annotation, uh, provided by, uh, business, uh, process outsourcing companies. Do you tap those? Do you subcontract? Do you acquire those to add to your roster? How do you interact with those people?

37:32Yeah, we typically have enough breadth in our database of people who are working from home to be able to deliver the needs of our customers. There are times when a customer will have data security requirements where it needs to be done in a facility and we support that model. And Appen does both. Appen is predominantly a work-from-home crowd-based model, but we also have many facilities around the world where we can support the more secure requirements of our customers. Yeah. And what's the, do you have kind of, I mean, of course, there's a bell curve distribution, but what is the top of the curve?

Read the full transcript

38:24what kind of people uh do you have doing this kind of work are they i don't know stay-at-home moms are they uh you know young people that the you know digital nomads or you know are they uh you know professionals that just want to be involved yeah it's a it's a it's a great question it's really fascinating so the mix of um i'll talk about appen previously which is um pre lln um it would be a mix of you know some of the things that you spoke about so stay-at-home parents people who are in remote areas and before work from home was such a big part of the market retirees and students, I would say, kind of the big categories.

39:21But as we get into the LLN space and the work that we're doing is increasingly domain-specific, I mean, we have people who have multiple PhDs who work for Appen who are just very passionate about the field that they work in and want to support the development of AI. So the market has shifted very quickly and now it's, you know, people who may have a full-time job and then they come home and can earn, you know, a really good hourly rate providing feedback into how the models are performing in an area of domain specificity that they're really passionate about. Can you talk about what the hourly rate is?

39:59I'm frankly thinking a few friends and relatives that might like that work. Yeah. So it, it does vary and it's, it's a significant variation and it can be, you know, the industry as a whole that happened in particular, you know, we pay above minimum wage every place we operate in. So there is a floor, right, and it does not go below that floor, right? So I think in the past, you know, maybe this is more mechanical Turk days where it might have been, you know, you get a penny for a task and you look back after an hour and you're like, well, that wasn't worth it. That's not the case anymore, which is great.

40:41And Appen's never been in that part of the market. So anywhere from kind of above minimum wage to hundreds of dollars an hour for a domain-specific expert. Yeah, yeah, that's interesting. Yeah, and so this is a growing space. do you think that the role of AI evaluation will crowd out the human evaluation? I mean, you were saying there will always be a role for humans, but how do you see that developing? I mean, also models are through various strategies are getting increasingly accurate. I see the market in two sides at the moment. So there is a B2C market, which is focused on consumer users, and a B2B market, which is focused on, you know, that's where you get into agentic AI and, you know, how our enterprises will be adopting AI to improve their operations.

41:48um the b2c market um i think will be dominated by a few companies that have very large distribution um in digital services so you've got google meta open ai in the west handful of companies in china and i think that's going to be the landscape there um and the interesting thing is that they they operate all globally. So I don't think we'll see a regional B2C sovereign models that actually have impact purely driven by the distribution. The other phenomenon that I see will play out is we've been as consumers as you know from the consumer side trained to exchange information about ourselves for some utility in a digital service.

42:45All right. All right. And this is a very well-trodden path. And that information about ourselves is used in advertising. All right. So there is a very well-established ecosystem in place to create that transfer of wealth. and I don't think we're there at the LLM point yet where this has been a major part of it, but you start to see what Google are doing and I think we'll see the similar thing with Meta and the interesting open AI at what point, if they do start to bring in advertising as a way to kind of offset the cost of the development and the inference for the models because I can't see a world where the subscription fees are going to offset the compute required to run these models.

43:42And in a bit of a long arc here, but I think what's going to be really important, and again, something that we've experienced through the market, is getting the user experience, that marginal part better, and getting the ad targeting marginal part better just has huge economic ramifications. So I expect that as AI becomes more pervasive than what we've had in the past with search and other forms of consumer applications, that there's going to be an increasing need for the subjective feedback into how these models are performing. And that will rise because of the economics, you know, the order of magnitude of the economics that support that environment.

44:36So you have a view across many, many different models. Are you generally impressed with model performance or are you kind of privately appalled? And thank God we have human evaluators. It's a little bit of both. there are areas where they work incredibly well and there are areas where uh you know i don't think they are living up to the hype so you know i'm excited about our role and you know helping and the way that the industry is progressing and i think it's the the pace of innovation is faster you know much but if you compare it to the you know you brought up feifei lee and image net that was incredibly fast paced this is i think 10 times faster so we're we're just moving at a pace that we haven't seen before so it's um it's it's hard to make an assessment at a point in time because i'll be wrong in two months from now and i've been asking people this what do you think has to happen right now it's it feels that we've hit kind of a plateau on on uh llm development i mean there's incremental improvement on this model or that um but there isn't the kind of breakthrough that came with the transformer introduction of the transformers transformer algorithm to me the the next step is uh solving uh catastrophic forgetting so that you can have a model learn continuously without erasing uh its its earlier training uh weights is what do you think is the next horizon that that the the industry is has got to get through look i think um when we when we cast our eye back over time right um there are three things that drive model performance it's some type of shift in architecture.

46:49It's scaling compute and scaling data, right? And, you know, I draw this curve where it's like the architecture gives a step change, like what we saw with transformers, like what we saw with, you know, convolutional neural networks. You had a step change and then this gradual improvement over time, right? And that improvement over time was driven by less about architecture breakthrough that's refining the architecture, but it's just throwing a hell of a lot of data at it and a hell of a lot of compute. So I think we're at this stage now, and you see it with the subsequent releases of OpenAI's models, the difference between GPT-2 to chat GPT-2 was enormous, right?

47:37And that was driven largely by human feedback, and then 3 to 4 and 4 to 5 less so, right?

47:47the breakthroughs that are needed to make this adopt at scale, it will be interesting whether there is an architecture breakthrough that creates this or it's just going to be the long slog of data and compute. And I think when you look at the investments that are being made in the industry at the moment, compute is clearly something that people are betting on. So it's, and it's a hard one because the model architecture breakthrough is really difficult to predict, almost impossible to predict. At what point will someone come with a new architecture? It's like trying to predict, you know, innovation.

48:34It's a nonlinear path, which makes it really exciting, right? Because there will be a point in time where someone comes up with either a slightly different architecture or a radically different architecture that has a huge breakthrough. But I think we're in this period of time at the moment where it's grinding out the improvements through scaling compute and scaling data. Is there anything I haven't asked that you'd like to talk about? I think I spoke a bit about the B2C side of the market. I think the B2B side is a really interesting one also. So, you know, as we think about evals and data curation and training, the B2B market is quickly evolving as a quite separate stream where, you know, the training data for agentic systems requires, you know, what's called reinforcement gyms, where people can go in and create a simulated environment with lots of different software, synthetic data, synthetic documents.

49:44The example is if you want to build an agent that is really good at creating financial documents in the oil and gas industry, you need to create an environment that is going to be able to go and train on something that looks like that. So we're seeing increasing demand for these highly demographic, highly industry and functional specific training environments. And then I think the next phase of that is making sure that the evals of that data is really accurate. So there's a part of the annotation market, which is can you find me a functional and domain expert, so a finance person with oil and gas industry, to go and evaluate how these molds are performing.

50:28So it's a really interesting market. I think the B2C side, there's a really clear path for massive human evals at scale to bring that global subjectivity and context. In the enterprise space and the B2B side, highly fragmented, small and narrow applications that are going to be built out. So it's going to be super interesting how both of these sides evolve, but I see them kind of quickly diverging.

From the publisher

This episode is sponsored by AGNTCY. Unlock agents at scale with an open Internet of Agents. 

Visit https://agntcy.org/ and add your support.


How do the world's most powerful AI models get trained and trusted at scale, and what does that really take from data to deployment?

In this episode, Appen CEO Ryan Kolln joins Eye on AI to unpack how rigorous human evaluation, culturally aware data, and model-based judges come together to raise real-world performance.

In this episode of Eye on AI, host Craig Smith speaks with Ryan Kolln, CEO of Appen, about building evaluation systems that go beyond static benchmarks to measure usefulness, safety, and reliability in production. They explore how human raters and AI evaluators work in tandem, why localization matters across regions and domains, and how quality controls keep feedback signals trustworthy for training and post-training.

Ryan explains how evaluation feeds reinforcement strategies, where rubric-driven human judgments inform reward models, and how enterprises can stand up secure workflows for sensitive use cases. He also discusses emerging needs around sovereign models, domain-specific testing, and the shift from general chat to agentic workflows that operate inside real business systems.

Learn how leading teams design human-in-the-loop evaluation, when to route judgments from models back to expert reviewers, how to capture cultural nuance without losing universal guardrails, and how to build an evaluation stack that scales from early prototypes to production AI.


Stay Updated:
Craig Smith on X: https://x.com/craigss 
Eye on A.I. on X: https://x.com/EyeOn_AI 

More from Eye On A.I.

All 266 episodes
#298 Ryan Kolln: How Appen Trains the World's Most Powerful AI ModelsEye On A.I. · 51 min
Listen in VO