In short
Eye On A.I. Podcast Episode Summary
Episode Title
#325 Phelim Brady: Why AI's Future Depends on Human Judgement
Host: Craig S. Smith Guest: Phelim Brady, Co-founder and CEO of Prolific
Episode Overview In this episode, Craig Smith engages with Phelim Brady to discuss the essential role of human judgment in the development and evaluation of AI systems. They shed light on Prolific, a platform that connects researchers and AI labs with real people for evaluating AI systems, emphasizing the importance of human input in a landscape that often appears fully automated.
---
Key Themes and Discussions
- The Human Layer in AI Development
- Human Involvement: Despite the perception of AI as fully automated, human evaluators significantly shape AI systems.
- Dirty Secret of AI: Brady refers to the underlying human involvement in AI processes as a 'dirty little secret,' highlighting how critical human input is for effective AI performance.
- Founding of Prolific
- Background of Prolific: Phelim Brady discusses his journey while pursuing his PhD at Oxford, realizing the need for high-quality human subject experiments online, leading to the creation of Prolific.
- Core Pain Points: Poor data quality, verification issues, and user experience were identified as significant challenges in existing platforms.
- Transition from Basic Tasks to Complex Evaluations
- Evolution of Work: Traditional platforms like Mechanical Turk facilitated simple tasks, but the complexity of data collection has grown, necessitating a focus on the skills and backgrounds of participants.
- Demographic Representation: Prolific emphasizes the importance of demographic awareness in model evaluation to ensure diverse perspectives.
- Model Evaluation and Benchmarks
- Shift from Traditional Benchmarks: The conversation highlights a growing skepticism towards traditional benchmarks as they may not accurately reflect real-world performance.
- Human Evaluation: There is an increasing reliance on real human feedback to provide a nuanced assessment of AI models rather than solely using standardized tests.
- Continuous Evaluation and Collaboration
- Human and AI Collaboration: Brady discusses how the relationship between humans and AI is evolving and why human judgment is indispensable even as AI systems improve.
- Long-term Role of Humans: The episode asserts that humans will remain integral to the evaluation process, especially in complex, subjective scenarios.
---
Notable Highlights
- Prolific's Unique Position: Prolific distinguishes itself by providing a vetted participant pool for nuanced tasks, differing from platforms that offer low-bid, commoditized labor.
- Scale of Operations: Prolific boasts a couple of million participants, with hundreds of thousands active monthly, leveraging word-of-mouth and referrals for growth.
- Diverse Applications: The platform supports a variety of projects, from academic research to AI evaluations, emphasizing its flexibility and adaptability.
---
Key Takeaways
- Human Judgment is Crucial: Effective AI systems rely on the insights and evaluations provided by real people, which are fundamental for the models to perform in real-world scenarios.
- Increasing Demand for Human Evaluation: As AI capabilities grow, the need for rigorous evaluation methods that incorporate human feedback becomes more prominent.
- Future of AI Development: The collaboration between humans and AI will shape future technologies, emphasizing a balanced approach to harnessing both human expertise and machine capabilities.
---
Conclusion This episode of Eye On A.I. serves as a reminder of the indispensable role of human judgment in the AI landscape. As the technology continues to evolve, integrating human insights into the evaluation process will remain a cornerstone of effective AI development.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIdentifying Pain Points in Human Data Collection
0:45 to 2:29
Discover the challenges in running high-quality human experiments for AI development.
“I started prolific while I was doing a PhD at the University of Oxford.”
The Role of Humans in AI Development
2:29 to 3:14
Understand the crucial role that human evaluators and labelers play in AI training.
“This is a pain point across research applications, so user research, polling, academic research, which was kind of our origin story.”
The Origins of ImageNet and Mechanical Turk
3:14 to 4:49
Learn about the historical context of AI datasets and their reliance on human labeling.
“kind of the dirty little secret of AI that it's built with humans, or that a lot of the other, these armies of human evaluators or labelers out there.”
Prolific's Approach to High-Quality Data
4:49 to 7:36
Explore how Prolific improves data quality through vetted participation.
“Like you have AI, but under the table, there are all these humans working at it.”
Vetting Participants for Reliable Data
7:36 to 10:00
Learn about the processes Prolific uses to vet its participants for accuracy.
“How much of your business is working with AI models and how much is, because I know you provide manpower or research, manpower I should say, for academic researchers.”
Recruitment Strategies for Diverse Talent
10:00 to 14:00
Discover how Prolific recruits a diverse range of participants for its platform.
“And a part of that investment and what you're selling is a vetted workforce, right?”
Understanding Platform Demographics
14:00 to 14:33
Learn about the talent distribution and demographics on the AI platform.
“And if I hear about it, I mean, you mentioned PhD students, but, you know, you certainly don't have a million PhD students on the platform.”
Onboarding Process for Contributors
14:33 to 17:06
Explore how contributors can join the AI platform and what is required.
“And can you give a breakdown like we've got, you know, 80 % are, you know, college educated doing simple tasks and 20 % are PhD students providing more difficult responses?”
Task Selection and Project Engagement
17:06 to 18:26
Discover how contributors select tasks and engage in projects on the platform.
“Or do you reach out when you get a query from a customer that wants a certain profile person working on their project that you pull together that cohort and then send that cohort an offer?”
Diverse Project Examples on the Platform
18:26 to 19:56
Learn about the variety of projects facilitated by the AI platform.
“What are the range of projects that people use the platform for?”
Show all 25 chapters
Evaluating AI Model Performance
19:56 to 22:44
Understand the methodologies used to evaluate AI model performance across different demographics.
“Some types of projects we optimize for maybe have a slightly higher complexity bar and a higher bar of methodological rigor.”
Persuasion Project Methodology
22:44 to 24:32
Delve into the experimental workflow of the AI persuasion project.
“develops, I mean, these different regions or countries are developing models that are trained on local data so they're more culturally attuned.”
Evolving Focus from Labeling to Evaluation
24:32 to 26:32
Learn about the shift from data labeling to model evaluation in AI development.
“But yeah, it was certainly more sophisticated than a just read the output of a model.”
Market Trends in AI Model Evaluation
26:32 to 28:00
Explore the growing demand for human evaluation in AI model performance.
“I mean, you know, in the early days of supervised learning, the focus was on labeling, particularly for computer vision.”
The Shift from Benchmarks to Human Evaluation
28:00 to 29:24
Learn about the transition in AI evaluation from academic benchmarks to human assessments.
“You can train your model to the benchmark or, you know, the benchmark somehow ends up in the training data and all sorts of things.”
Evaluating AI Applications in Real-World Contexts
29:24 to 31:00
Discover how enterprises approach evaluations for their AI applications.
“I heard million, was it a million or multiple millions or something on the platform?”
Challenges in Model Selection for Enterprises
31:00 to 32:30
Understand the complexities enterprises face when selecting AI models for specific applications.
“I mean, there's evaluation of the frontier models.”
The Role of Human Judgment in AI Evaluation
32:30 to 34:14
Explore why human judgment remains critical in evaluating AI despite advancements in models.
“Yeah, so enterprises, I mean, if I'm an enterprise and I'm building an application and I have a choice of, you know, eight different models to hit for my inference, you know, through an API.”
Future of Human and AI Collaboration
34:14 to 36:24
Gain insights into the evolving relationship between humans and AI systems in evaluations.
“And then also understanding the delta between the performance and outcome that they're driving towards for their application and then the gap between that and their perhaps fine-tuned model or customized model.”
Leveraging AI for Prolific's Operations
36:24 to 38:16
Learn how Prolific integrates AI into their platform and operations for better efficiency.
“model confidence is low or where you want to cherry pick and just make sure that the models and evaluators are performing as expected.”
Business Model and Payment Structures of Prolific
38:16 to 41:20
Discover how Prolific's business model works, including payment structures for users.
“I mean, either in, you know, managing the platform or sourcing participants or, yeah, I mean, is there AI in your platform?”
Using Prolific for Surveys and Polling
41:20 to 42:04
Understand how Prolific can be utilized for running surveys and engaging participants.
“And it's purely just the amount of data that you collect and the amount of time that you ask from the contributors.”
The Role of Prolific in Data Collection
42:04 to 43:31
Learn how Prolific supports researchers in complex data collection processes.
“If you're in the business of collecting data or running research, typically you have many projects, or even if you move companies or universities, bring your preferred tools with you.”
Future Aspirations for Prolific
43:31 to 45:16
Discover Prolific's vision to become a comprehensive human data platform.
“And yes, all of the projects or the pay is based on some combination of the length, complexity of the project and complexity of the requirements on the audience side.”
Challenges in Humanoid Robot Data Collection
45:16 to 47:16
Understand the challenges of data collection for humanoid robots and new explorations.
“And yeah, much more work to be done on all of those strands.”
Transcript
Automatic transcript. May contain errors.0:00Phelim Brady:I have in past interviews with people working with human resources for AI is that it's kind of the dirty little secret of AI that it's built with humans or that a lot of the other these armies of human evaluators or labelers out there.
0:22Craig Smith:or computational biology, but while there, realized that there was core pain points in running high-quality, academically rigorous human subjects experiments online. My name is Phelan Bradley. I'm the co-founder and CEO of Prolific, and Prolific is a human data platform. So we connect both the infrastructure methodology and high quality global pool of humans and connect them with data collectors for purposes of research and human data for AI, in particular post-training and evaluation use cases. I started prolific while I was doing a PhD at the University of Oxford. I actually did my PhD in bioinformatics or computational biology.
1:14Craig Smith:But while there, I realized that there was core pain points in running high quality, academically rigorous human subjects experiments online. So platforms that were popular at the time had problems of poor quality data and verification of the humans who were taking part in research. The tooling and infrastructure was extremely poor user experience for both sides of the platform. again impacting both the ease of use and the quality of the resulting data. And ultimately, it created an unhealthy platform and dynamic. We set out to build Prolific in order to solve that problem. And originally, it was a side project while I was completing my PhD and grew the company to reasonable traction before completing my PhD.
2:13Craig Smith:And the story since then really has been rediscovering that same core pain point across a range of different industries. How do I recruit high quality data online from an audience of highly verified trustworthy participants? This is a pain point across research applications, so user research, polling, academic research, which was kind of our origin story. And now more recently, also human data in the AI development lifecycle. And I think we try to apply some of this methodological and academic rigor of the behavioral sciences to some of the problems in evaluation of generative AI applications.
2:54Yeah.
2:55Phelim Brady:And this is an interesting area. And I know some of your other interviewers have pointed this out, as I have in past interviews with people working with human resources for AI, is that it's kind of the dirty little secret of AI that it's built with humans, or that a lot of the other, these armies of human evaluators or labelers out there. My experience or my knowledge of this goes back to interviewing Fei Li, who I just had on the program. But the first time I interviewed her was about ImageNet, which was this massive database, really the first image database for training supervised learning programs.
3:50Phelim Brady:And it was her data set that allowed Jeff Hinton to validate deep learning in 2012 with Alex Nutt and, you know, kick off the deep learning revolution. And she turned to Mechanical Turk to do the labeling of her images. At that point, I hadn't heard of Mechanical Turk, but it's kind of funny because it's a reference to a 17th or 18th century, I guess, 18th century German automaton mannequin that played chess. And famous people came and played chess with it. I think Ben Franklin played chess to it, with it, and it would win, and people were just amazed. What they didn't know was there was a small statured chess master curled up in the cabinet underneath the chessboard who was working the mannequin with levers and things.
4:48Phelim Brady:So it's kind of like that. Like you have AI, but under the table, there are all these humans working at it. So that was Mechanical Turk. And then I worked for a long time and had on the podcast guys who had a platform called Labelbox. Are you familiar with them?
5:07Craig Smith:Yes, they're Labelbox.
5:10Phelim Brady:Yeah, their labeling platform. And their thesis sounded similar to yours, that there isn't a good platform. They built this platform, and you can plug in third-party BPO teams into it, or you can have your own labelers. But it's a platform that they don't manage the human resources, or they didn't at that time. And then, as you mentioned, I spoke recently to Appen, the field's teams, or rather coordinates teams to do reinforcement learning with human feedback on foundation models. So where do you fit in that tradition besides being the chess master underneath the chessboard?
5:58Craig Smith:that's a great great question a great i think framing of of of the space i think we fit most in the tradition of the mechanical turk so trying to abstract away the complexity of dealing with the messiness of real humans in the real world and provide labs and data collectors the kind of high quality data that they want. I think in some ways you could describe Pritifik as what could and should have become. And I think a key kind of shift in the market, which we've invested heavily in, I think in contrast to a platform like Mechanical Turk is Mechanical Turk was great at applications like ImageNet that you mentioned, where, you know, labeling cats in images, is this a hot dog style questions any human would do.
6:59Craig Smith:So it's a fairly commoditized task. And as a result, every human on the other end who's doing that labeling is somewhat fungible or replaceable with any other person. That's changed quite a lot since the days of ImageNet. And now the audience of the people who are annotating your data or providing your post-training data, providing your RLHF data really matter, their background, their expertise. And I think coming from behavioral research context, the representation or representativeness or the generalizability of the audience, that I'd say is like where we have focused to date in providing this human intelligence layer, but also maximizing the breadth and depth of the audience that you're able to access through a platform with an aspiration that the best and depth of humanity is ultimately reflected and encoded in these AI models.
7:55Yeah.
7:56Phelim Brady:How much of your business is working with AI models and how much is, because I know you provide manpower or research, manpower I should say, for academic researchers. How much of it
8:12Craig Smith:is working with a models roughly 50 50 and in terms of the business in terms of the investment and manpower we put behind the platform although i think crucially there we focus on areas where there is strong synergy between the the two sides of the platform so the audience choice and the representativeness and the ability to tap into real world participants in order to understand and real-world behavior is a shared requirement and kind of pain point across both industries, as well as the quality of the data collection tooling and the infrastructure that we provide. I think there's two kind of core dimensions into data quality.
8:52Craig Smith:It's the quality of the contributors or the participants who are providing that data and then quality of the infrastructure and the methodology that you use in order to analyze and abstract that data. So those two components are core investments that we think that provides synergy across the markets that we operate in. And then obviously there's going to be differences in terms of like how we serve those customers, whether they're large enterprise and need deep integration into their on-prem tools, or they're a PhD student who needs easy, flexible, self-serve access. And I think crucially, fundamentally the frontier of AI capabilities and AI research is a research problem.
9:38Craig Smith:So there's a strong overlap between academic research universities and the talent who's driving forward the capability of these labs. So our purpose as a company really is to accelerate the frontier of human-centered or transformative research and AI. And there's a lot of kind of shared investment that supports both of those challenges.
10:00Phelim Brady:Yeah. And a part of that investment and what you're selling is a vetted workforce, right? I mean, this isn't like Mechanical Turk, which has very little vetting, as I understand. I've never used it, but it's kind of self-serve. You sign up and pick a task and execute it. And consequently, the tasks are very simple. But you're further up the food chain, dealing with more nuanced tasks. So you vet your workforce. How do you vet them?
10:38Craig Smith:First, I may be slightly correct the term workforce in that our participants are folks who participate in Prolific in a flexible manner. And as a result, we're able to reflect kind of real world users, either consumer samples or people who have full time jobs in other professions and are able to participate in Prolific in an ad hoc manner. So slightly different from the contractor-style workforce of some of the other platforms that you mentioned. And the verification and the vetting side of things is one of our core investments and core IP. We have many layers of protection from the fairly obvious things in terms of doing identity verification of all of the participants.
11:22Craig Smith:We know who they are. Not only do we know who they are when they sign up, but we check this on a periodic basis. So we know that person is still the same person. deep profile information to make it easy to route the right tasks at a right person based on either kind of independently verified data points or qualification style gates that we put folks through in order to run them through an exam as well as lots of behavioral analysis and understanding like how these folks are actually interacting within the data collection or experimental workflows. But I think the latter is increasingly crucial in the world of agentic fraud.
12:04Craig Smith:So people being able, agents or AI systems being able to replicate the behavior of humans very accurately is kind of a new threat to online data collection and primary data collection more generally. And one that we feel like we're right at the cutting edge of being able to defend against while still providing the scale of audience choice and global participation that we're able to power. Yeah.
12:30Phelim Brady:And so what would be, first of all, how many participants do you have on the platform? Not people using the workers, but the people providing the human talent. How many of them are there?
12:47Craig Smith:Yeah, so we have a couple of million people on the platform in general with typically several hundred thousand active in any given month.
12:57Phelim Brady:Wow, that's impressive. And I mean, obviously you're not going through LinkedIn and reaching out to individuals if you've got, you know, a million people. How are you reaching them?
13:12Craig Smith:Yeah, it's a combination of factors. the core growth of the platform comes through word of mouth growth and referrals. So because we offer a positive experience and meaningful work and meaningful pay, this is something that people tend to share online, share on social media, share with their family and friends. That's a core growth engine for us. Secondly, complement that with some incentivized referrals. So where we have gaps in our audience that our customers want to tap into, let's say we're looking for a PhD in a particular aspect of biology, we typically will have a small number in our pool and we can use that to bootstrap a larger sample through a folks network.
13:57Craig Smith:And then third, we do community and event-based marketing in order to supplement the first two channels. But I'd say our core growth and our core engine on both sides of the platform is the trust and integrity, high-quality experience that we're able to offer, leading people to share this amongst their colleagues or their friends. They're friends.
14:21Phelim Brady:Yeah. And if I hear about it, I mean, you mentioned PhD students, but, you know, you certainly don't have a million PhD students on the platform. What's the range of sort of talent level? And can you give a breakdown like we've got, you know, 80 % are, you know, college educated doing simple tasks and 20 % are PhD students providing more difficult responses?
14:56Craig Smith:I don't have the precise numbers to the hand, but I would say we generally optimize the overall platform to reflect the dimensions that exist within real world populations. So try to reflect the demographic breakdown that exists within a representative US sample or representative UK sample, etc. And then in terms of the AI work specifically, I'd say it's maybe even roughly a third between general audience sampling. So this is consumer-based testing where the representativeness and the generalizability of the audience is key, but no particular specialism or skills are required. And then we have tasker-based work where we either qualify or train people within the crowd to be high-taste evaluators or people who have some skills in general AI evaluation and data labeling, which requires a different level of attention and understanding of how these models work.
15:59Craig Smith:And then the final third would be expert level workflows where the subject matter expertise that folks bring from their prior experience is the critical dimension of whether they are appropriate for the project.
16:12Phelim Brady:Yeah. And so if I wanted to make some money on your platform, what would I do? What's the onboarding process?
16:22Craig Smith:Yeah, to be onboarding process from a contributor or participant view is pretty straightforward. So it's a kind of a five-step process in terms of filling in some basic information, giving us a short interview on your background and experiences and where you think you might have unique skills or experience that would be valuable either to researchers or to the AI labs. and then some verification of both your backgrounds, this KYC style object, know your customer, as well as a short kind of behavioral assessment to validate that you're engaged, attentive, and trustworthy. Yeah.
17:04Phelim Brady:And then I'm on. And then is it a list of tasks that I can choose from? Or do you reach out when you get a query from a customer that wants a certain profile person working on their project that you pull together that cohort and then send that cohort an offer? I mean, self-serve or?
17:28Craig Smith:yeah yeah it's sort of between kind of one-off projects or tasks which you'll be notified for on your project on your path on your dashboard or longer running projects where you need to agree to a certain level of commitment up front maybe these are going to last for several weeks or months or require a little bit more time involvement where perhaps there's going to be an initial assessment and then if you get past that initial assessment you're able to participate in the longer running project. Yeah, it's kind of a double opt-in marketplace. So every task is advertised to you. You're able to assess whether or not that's something you're interested in at the reward that is being offered.
18:11Craig Smith:And you trust that you understand how your data is being used. If you're interested, you can then accept. If not, you reject the place and that opens up then to another participant on the platform to choose if they want to take part. Yeah.
18:27Phelim Brady:What are the range of projects that people use the platform for? I mean, again, you know, you've got PhD students that are doing research projects. What kinds of, can you give us an example of that? And then on the low end, I guess the low end would be data labeling or in what kinds of projects do you use? Can you give us an example there?
18:56Craig Smith:Yes. So one of the beauties of the platform, I think, why platforms like Mechanical Turf became popular in the first place also is the flexibility of use case. So we've run nearly half a million projects over the last year with a vast range of different requirements and outcomes resulting in thousands of publications as well as like many capability improvements to models that don't get as well publicized or cited. But maybe to give you a few recent examples, we actually don't do that much work on the, as you've described, the low end of the data labeling. As I mentioned at the start of the conversation, we focus on high quality projects where the audience and who is behind your data is a critical component.
19:46Craig Smith:And most of the more commoditized data labeling use cases don't have that requirement, maybe a better fit for a platform that optimizes for lower cost of labor. Some types of projects we optimize for maybe have a slightly higher complexity bar and a higher bar of methodological rigor. So we've run a project recently which was investigating the persuasive capabilities of state-of-the-art models, which was run by the AI Security Institute in the UK. And this is really this, I think it's a nice example as a combination of the behavioral research skills that we bring and the capabilities we bring in model evaluation.
20:22Craig Smith:So really understanding from a safety perspective, how politically persuasive can these AI models be in the real world with a representative selection of an audience. Another example is a project that we've run ourselves called Humane, which is a user preference model evaluation benchmark. So there are popular platforms like LLM Arena or Chatbot Arena, where in short, two state-of-the-art models are pitted against each other and users need to select which model they prefer and why. And ultimately, the preference of these models, these participants are analyzed and this results in a leaderboard of which models are preferred by users.
21:11Craig Smith:One of the challenges with Chatbot Arena is there is no control on the audience or the people who provides these judgments or preferences, which I think is analogous to running a political poll without controlling for the audience selection. So we've run some projects which take the same idea, battle Model A against Model B, but add some methodological rigor in terms of the demographic representativeness of the audience behind the leaderboard. All of the models are double blinded, which essentially then allows you to ask the question of all of these models, is there a difference between the preferences of a US population, a UK population between left-leaning participants or right-leaning participants?
21:59Craig Smith:Is age a factor in how people rank the performance of these models, in which context do people prefer Model A or Model B, adds a lot more nuance to what otherwise is a fairly simplistic leaderboard. And the short version of the results of that project is that the ranking of models does change based on the demographics and the audience behind the models, which I think is going to be an increasingly important factor as these models are rolled out in the real world, in different geographic locales, people with different culture are using them for different use cases, understanding which model is best for which context is going to be a more interesting question perhaps than which model is objectively state-of-the-art on an overall basis.
Read the full transcript
22:43Phelim Brady:Yeah, particularly as sovereign AI develops, I mean, these different regions or countries are developing models that are trained on local data so they're more culturally attuned. But I'm interested in the persuasiveness project. How did that work? I mean, can you describe that? I mean, you have people in your database and they talk to a model or they read outputs from the model or what's happening there?
23:22Craig Smith:Yes, there were many manufacturers there and it was an experimental workflow. So essentially, people were split into randomized conditions. And as a result, there was kind of an AD style test of whether participants, if they were exposed to a back and forth conversation with one of, I think, about 20 models, and these models were instructed to persuade the participants to agree with one of political stances using different rhetorical strategies. So the experimental methodology was fairly sophisticated. And then the participants were measured on whether they agreed with the issue before they were exposed to the model and the conversation and after, and then as a result, you're able to measure the difference in how different models and different different strategies informed how much these participants changed their minds on the topics that were at hand.
24:25Phelim Brady:Yeah, but the interaction between the participant and the model, is it like, you know, have a conversation with this model for 30 seconds and, you know, about some prompt, some topic or is it that there is a prompt and the model responds and the participant reads the response or is the participant actively engaging in a conversation with the model
24:58Craig Smith:it was i believe it was the latter so an active engagement multi-turn multi-turn conversation i'm not sure if the i don't have the exact constraints to mind i think it was perhaps a minimum number of interactions that were required in order to accept the data point. But yeah, it was certainly more sophisticated than a just read the output of a model. It was a kind of a live interactive style of experiment. Yeah.
25:22Phelim Brady:But that's a really interesting use case. So you need a certain level of political awareness. I mean, how did you pull together that cohort?
25:37Craig Smith:yeah in this case this i think goes back to the representation and the generalizability of the of the audience so trying to get a sample of of the real world so that the the outcomes of that experiment are more likely to generalize to real world scenarios and to a real world context was a crucial criteria of of that that kind of project and we have the tooling built into the platform which allows you to do census-matched sampling, which means that the resulting participants you get broadly match the population on key criteria, such as age, ethnicity, political affiliation, things of that nature.
26:19Yeah.
26:20Phelim Brady:I mean, you said you don't do much sort of labeling anymore. Is that right?
26:26Craig Smith:That's right. We never did a huge amount of, let's say, the simple kind of commoditized style of labeling. One could call the post-training evaluation data sets a form of data labeling.
26:37Phelim Brady:Yeah, yeah. Actually, that's what I was getting at. I mean, you know, in the early days of supervised learning, the focus was on labeling, particularly for computer vision. And then just in the last year, as these models have matured and become more widely used, and as there are more models competing, evaluation has sort of come to the fore. How long have you been doing model evaluation?
27:09Craig Smith:Yeah, we've been working in that area, I would say, for two to three years. So I think post the chat GPT boom, the shift of user-driven model evaluation has shifted towards, as I said, kind of more complex audience requirements, more expert participants and a requirement for more kind of scientific rigor. And that's been a natural fit for our platform capabilities, which was the original wedge within this space. We've obviously invested pretty heavily in additional capabilities and frameworks and the ability to kind of convert human judgment into useful signal for these model developers.
27:55Phelim Brady:Yeah. Yeah. And that's, yeah, as the models, I mean, evaluation on benchmarks has gotten sort of a, people have soured on that to a certain extent because you can, there are all sorts of problems. You can train your model to the benchmark or, you know, the benchmark somehow ends up in the training data and all sorts of things. So people don't trust, you know, a model winning the math Olympiad anymore. That at one time was a big deal. Now they're more focused on this human evaluation. does that is the market like blowing up for you guys that end of the work yeah i think that is a
28:42Craig Smith:key driver of our expansion in the market this shift away from like the academic benchmarks or exam-based benchmarks which the frontier labs use to kind of advertise their relative performance there's such a strong incentive as you said to to either to unintentionally gain these benchmarks and leaderboards so their usefulness is decreasing as they become saturated and essentially the models are able to all pass with with flying colors but then it's becoming harder to judge the the real world evaluation and i think that's where platforms like ours come to their own and in trying to simulate the real world environments in which these models and agents are interacting with real people trying to achieve real problems and then evaluating that performance is ultimately the real the real goal that's ultimately going to drive kind of the the economic impact of
29:40Phelim Brady:these models have in in practice yeah you guys you know are part of an industry you say you have I've forgotten the number. I heard million, was it a million or multiple millions or something on the platform? On the evaluation side, if you were to look globally, do you have a sense of how many people are involved in evaluation through platforms like yours?
30:11Craig Smith:I don't have a strong estimate of the global number of people involved in this type of work though I would say it is it's increasing and also the the types of people who are likely to be involved in it over the next couple of years are likely to shift considerably as the labs kind of shift their focus towards knowledge work and understanding the capabilities that these models have in kind of all of the areas of the economically productive labor, folks within all of those industries are going to have some involvement in evaluating and providing feedback to the tools that they use day in, day out in order to start to improve their performance, not only on these academic benchmarks, but actually how they're performing in real world conditions on real world problems.
30:59Phelim Brady:Yeah. I mean, there's evaluation of the frontier models. That's one thing. But then there are all these applications built on top. Is it companies building applications that come to Prolific to evaluate their applications? Or are there enterprises that are adopting solutions who want to evaluate? Increasingly, both.
31:26Craig Smith:I would say that the labs are leading the way as they have the talent in-house to understand the methodology in the frameworks that are required to build out benchmarks. And the majority of labs will have in-house teams and in-house evaluations that are accustomed to their requirements and are not necessarily always one-to-one mapping to the public benchmarks. But as enterprises start to build on top of this infrastructure and build agents for their specific use case or their specific context, they increasingly also need this evaluation, not only of their own tools, but also because there is such a wide choice now of which open source or frontier model to use, even evaluating which base model to use in the first instance is a useful piece of data and is not necessarily obvious up front.
32:14Craig Smith:But then as they build their application, in particular understanding, I think safety and reliability and trustworthiness in real-world conditions is often they have a higher bar for that kind of robustness in specific applications than perhaps the infrastructure layer is going to have.
32:29Phelim Brady:Yeah, yeah. Yeah, so enterprises, I mean, if I'm an enterprise and I'm building an application and I have a choice of, you know, eight different models to hit for my inference, you know, through an API. Yeah, you know, I've seen sort of casual rubrics where, you know, Claude is good for this, ChatGPT is good for that, and, you know, Gemini is good for something else. But, I mean, who knows on your specific use case. So are you having enterprises come to you and say, we need to figure out which model works best in our application and do A-B tests, as you say, or A-B-C-B tests or something.
33:21Craig Smith:Yeah, I mean, I think that's you precisely laid out one of the challenges that we want to tackle, which is many of this evaluation is based on vibes and based on intuition that domain experts have. We want to bring a layer of objectivity and rigor to this evaluation, which I think is going to be particularly important in enterprise applications and particularly in regulated or sensitive domains like healthcare, finance, law, things of that nature. And then on the second part of your question, that choice of base model is just the first step, right? I think there's many more steps than in terms of testing for safety and having a continuous flywheel of feedback, both from automated evaluations.
34:03Craig Smith:We've talked mostly about human judgment evaluations, but there's also a whole strand of technical and automated evaluations, which is a complement to the human judgment. And then also understanding the delta between the performance and outcome that they're driving towards for their application and then the gap between that and their perhaps fine-tuned model or customized model. And then critically, the tooling and the products to help them bridge that gap between the capability that they're after and the kind of baseline capability that they get with the existing model.
34:38Phelim Brady:You mentioned agents at the very beginning in the context of ensuring that you have a human on the other end doing the evaluation and not an agent. But the agents are being employed in evaluation and they're getting better. Is there a reason why your industry wouldn't eventually fade away as agents play that role?
35:06Craig Smith:Just to be clear, we're active users of agents and LLM as a judge and similar tools in order to augment human performance and make sure that we're using the relatively expensive component of human judgment in an intentional kind of cost effective way. I think there's a few reasons why I think at minimum it's going to be a long time before
35:31Phelim Brady:the human component dies out.
35:33Craig Smith:Firstly, I think the human judgment is where the alpha is. So if you're looking to push the capability beyond what is already available, almost by definition, you are not able to build an automated evaluator that is able to assess that gap in capability. So you need human judgment, human judgment there. Of course, in cases like chess, the models have gotten to a stage where the human judgment is no longer kind of a net benefit to improving the models. But I think chess and similar verifiable domains are a fairly small set of the useful tasks. And wherever there is kind of ambiguity or subjective opinion required, human judgment is going to be in the development lifecycle for a long, long time, as well as I think crucially the trust and safety component where I think at minimum having an escalation path where the model confidence is low or where you want to cherry pick and just make sure that the models and evaluators are performing as expected.
36:37Craig Smith:So certainly I think the human judgment and the human's role in this lifecycle will change. I think the relative value that it provides is going to increase, if that makes sense.
36:49Phelim Brady:Yeah, absolutely. Yeah, and I didn't ask that question as a challenge to your viability. I'm just imagining the future, but that is one of the really interesting things about AI. I mean, these large models is they encode human knowledge off of text, and there's a lot of ambiguity and a lot of misinformation included in the text. And then we'll ultimately need humans to guide the models. And it sounds, you know, you didn't come back with a number, but it sounds like there must be millions of people around the world engaged in this. Do you think that's an overstatement? Millions?
37:37Craig Smith:I don't think that's an overstatement, for sure and also i think the many more people are going to be involved at minimum in a passive way as more and more of these models are rolled out in as as products and i think just maybe touch on your other other point we've worked in with academics in the field of human computer interaction for for a while i think that field is shifting maybe towards like human ai interaction or human agent interaction so i think there's a lot of work to be done in actually understanding like how to get the best of both humans and AI systems, and how do you make the collaboration and the engagement between them better than the sum of its parts, as it were.
38:16Craig Smith:This is not necessarily kind of obvious how it's going to work at Apri, and I think will require quite a bit of intentional effort in order to understand how to develop systems that work well for humans and make sure these agents and systems are human-centered and are supportive of folks' work.
38:36Phelim Brady:How much AI do you use in a prolific? I mean, either in, you know, managing the platform or sourcing participants or, yeah, I mean, is there AI in your platform?
38:51Craig Smith:Yes, and an increasing amount. I think there's an interesting fact that AI is quite a convergent force in technology in the sense that increasingly, I think products are evolving to the magic AI boxes where you type in your request and you expect the AI to do its magic and return its result. I think the very early version of the we've rolled out is on the data collector side, the audience requirement tool. So you just express who you're looking for in natural language and we go away and we use the data that we have on our participants in order to find the right people. rather than you needing to really think about the and or requirements of how you select for those people.
39:33Craig Smith:And similarly on the contributor side, instead of asking them a battery of questions to understand who they are, we're able to offer them a path where they're able to interact in natural language through audio and video. And then we're able to use that richer qualitative data in order to extract what's important to the other side of the platform and improve that matching. and I think that abstraction will increase where the magic box will you'll ask it instead of find me these people it will be okay and design me the data collection tool and please go away and actually execute on the the work and then ultimately maybe automate the full workflow so the researchers or data collectors kind of stay in the outer loop of the project and designing the hypotheses and the outcomes that you want but increasingly the the agents and the AI tools that we build allow them to automate a lot of the details.
40:27Phelim Brady:Yeah. And you're doing that now or that's on the roadmap?
40:31Craig Smith:This is the roadmap. So the two examples I gave at the start of the define the right audience through natural language and the understanding of the contributor's background through AI interaction. Both of those are live products and I think early examples of more to come. Yeah.
40:48Phelim Brady:How does the customers, whether it's an individual PhD student or a foundation frontier model company, how do they pay? Is this a subscription or a pay as you go?
41:04Craig Smith:Yeah, it's a fully usage-based billing cycle. So we will recommend the recommended pay to the contributors, and then we take a portion of that payment. But it's fully usage-based. So you can go from tens or hundreds of dollars through to tens of millions. And it's purely just the amount of data that you collect and the amount of time that you ask from the contributors.
41:28Phelim Brady:and how much of your business is kind of retail in that sense that you know people are coming on the platform for a short project and then they're off and how much is it big enterprises that have longitudinal studies going on that they need you guys for yeah we span the spectrum
41:56Craig Smith:of budgets and wallet sizes and use cases. I'd say the vast majority of users have recurring use cases. If you're in the business of collecting data or running research, typically you have many projects, or even if you move companies or universities, bring your preferred tools with you. So yeah, we see the majority of customers come back on a repeated basis.
42:19Phelim Brady:Yeah, you mentioned polling also. Do people use you to run surveys?
42:28Craig Smith:Yes, again, this is a use case that is supported. And again, I think something that we're increasingly interested in as the complexities of running high quality polls is becoming more and more challenging for a variety of reasons. Harder to reach people, harder to incentivize people. So I think that's an interesting role that we can play there. Yeah.
42:53Phelim Brady:And on the compensation side, each project is the contributors are paid according to the project. So there isn't a standard hourly fee or task fee that they're earning. And is it enough that people, they can supplement their income certainly, but can people earn a living income as a contributor?
43:21Craig Smith:We don't optimize for folks who want to earn a living on the platform. So primarily it's a side hustle or supplementary income. And yes, all of the projects or the pay is based on some combination of the length, complexity of the project and complexity of the requirements on the audience side. But all of that is made transparent to both sides of the platform. People have that openness ability to kind of, it's not a black box in terms of us going away and not showing how the sausage is made. The data collectors and the participants are able to interact with each other and see those sides of the platform as well.
44:00Yeah.
44:01Phelim Brady:And where do you see Prolific going? I mean, as I said, I would guess that there's growing demand as more and more companies release these models. But then there's also, it seems to me there would be opportunities, for example, on polling, where not only do you source the respondents, but you could run analytics on top of those responses. I mean, if the data is flowing through your pipeline. So where do you see Prolific going?
44:40Craig Smith:Yeah, totally. I think it's exciting use cases across a range of applications. And ultimately, our ambition is to become a full stack human data platform. So provide a global, high quality participant pool where you're able to access both the breadth and depth of humanity. the tooling, whether that's through partnerships or through our own tool to tap into those humans for a wide range of different applications, and then supporting the methodology and the tooling really to advance the frontiers of research and AI. And yeah, much more work to be done on all of those strands. And yeah, many exciting opportunities, both in the developing transformative or human-centered AI, and then also really understanding how this is changing and influencing human behavior and really being the platform that provides this authentic, high-integrity human data in the age of lots of AI-generated data as well.
45:42Phelim Brady:Yeah, and I know I'm up past the hour. Do you have time for another question or two?
45:50Craig Smith:Unfortunately, I probably need to drop one more question. Absolutely.
45:57Phelim Brady:Yeah, there's last year or in the past year, there's been this boom in humanoid robots. But the critical block bottleneck for humanoid robots is data collection, you know, because with LLMs, you've got the internet full of text, but humanoid robots, you need, you know, data from humans manipulating objects or, you know, identifying things in scenes. Are you doing any work on data collection for humanides?
46:30Craig Smith:Yes, it's not the core of the work that's done on the platform to date, but is, I would say, an area of open discovery. In particular, you mentioned you had Faye Lynn on the podcast recently, and she probably talked quite a lot about world models being a required environment in order to accelerate the training of embodied AI. So we've actually done some very interesting work on integrating prolific into virtual environments so participants can take part in data collections in VR. So I think we've gone from kind of text, images, video. I think world models and virtual reality are an interesting modality that we'll definitely explore in the coming years.
47:14Phelim Brady:Okay.
From the publisher
AI often looks fully automated. But behind the scenes, a huge amount of human judgment is shaping how these systems actually work.
In this episode, Craig Smith speaks with Phelim Bradley, co-founder and CEO of Prolific, a platform that connects millions of real people with researchers and AI labs to evaluate and improve AI systems.
They explore the hidden human layer behind modern AI, why traditional benchmarks are becoming less reliable, and why AI companies increasingly rely on real human feedback to measure model performance in the real world.
Phelim also explains how demographic differences influence how models are evaluated, why human judgment remains critical even as AI improves, and how the collaboration between humans and AI will shape the next phase of development.
This conversation reveals the human backbone behind today's AI systems.
Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI
(00:00) Preview and Intro
(02:45) Founding Prolific And Early Pain Points
(06:30) From Mechanical Turk To Representativeness
(09:55) Academic Research And AI Use Cases Split
(13:40) Vetting Real Participants And Fighting Fraud
(17:45) Scale, Community Growth, And Talent Mix
(22:00) High-Complexity Projects Over Commoditised Labeling
(26:40) Measuring Model Persuasion With Live Conversations
(30:20) Demographic-Aware Model Preference Benchmarks
(34:10) The Rise Of Human Evaluation Over Benchmarks
(38:00) Enterprise Model Choice And Continuous Evaluation
(42:00) Why Humans Won't Disappear From The Loop




