In short
Practical AI Podcast Episode Summary
Episode Title
Suspicion Machines ⚙️
Podcast Overview
- Focus: Making artificial intelligence practical, productive, and accessible to a wide audience.
- Themes: Discussions on AI topics such as machine learning, deep learning, neural networks, and real-world applications of AI.
Episode Description In this episode, hosts Justin and Gabriel examine "suspicion machines," which are machine learning algorithms used in welfare systems across Europe to flag potential fraud cases. They share insights from their investigative journalism that involved analyzing these models to understand their implications and biases.
---
Key Concepts and Discussions
What Are Suspicion Machines?
- Definition: Algorithms that assign risk scores (ranging from 0 to 1) to welfare recipients, estimating their likelihood of committing fraud.
- Context: These systems have become controversial in the welfare debate across Europe, where political discussions revolve around fraud allegations and welfare distribution.
Investigation Findings
- Case Study: The Netherlands, where 30,000 families were wrongly accused of fraud due to a faulty machine learning model.
- Research Approach:
- Utilized freedom of information requests to access documents, manuals, and eventually the model code.
- Detailed assessment of model behavior demonstrated systemic biases and inaccuracies in fraud detection.
Challenges in Algorithmic Transparency
- Access Issues: Difficulty in obtaining complete model files due to claims that full disclosure could enable fraud.
- Literature Gap: A noted lack of research and reporting on European welfare systems' use of machine learning, compared to the extensive discussions around similar technologies in the U.S.
Technical Insights
- Model Characteristics: The investigated model used a gradient boosting machine algorithm and was based on a variety of demographic and behavioral features.
- Data Limitations: Issues with the dataset included the non-random selection of training data and a lack of clarity on true fraud versus unintentional error.
Bias and Fairness
- Feature Analysis: Some features in the model were deemed discriminatory or problematic (e.g., assessing appearance, language skills).
- Disparate Impact: The model appears to disproportionately flag certain demographic groups as higher risk, raising questions about the fairness of these assessments.
System Performance
- Hit Rate: The model had a 30% success rate in identifying fraud, which, while better than random guessing, still indicates significant flaws in the system.
- Selection Bias: Concerns that training data was influenced by the methods of selection, leading to skewed results.
---
Key Takeaways
- Critical Assessment of AI: The discussion emphasized the need for deeper evaluation and transparency regarding machine learning applications in sensitive domains like welfare.
- Impact on Individuals: Highlighted the punitive measures faced by individuals flagged by these systems, which can lead to severe consequences regardless of actual guilt.
- Future Implications: The episode encourages ongoing discussions about the ethical use of AI in public policy, suggesting that practitioners consider broader societal impacts rather than just technical performance.
Thoughts on the Future
- Transparency and Accountability: Both hosts advocate for more discussions on algorithm transparency to prevent harm and ensure fair treatment in welfare systems.
- Societal Perspective: The need for a holistic view of AI systems that balances technical efficacy with social responsibility.
---
Conclusion This episode of Practical AI shines a spotlight on the ethical implications and practical challenges of deploying machine learning in welfare systems. It calls for greater scrutiny, improved transparency, and a focus on societal implications of AI technologies in public governance.
---
References and Further Reading
- Article: [Inside the suspicion machine](https://www.wired.com/story/welfare-state-algorithms/)
- Methodology Behind the Investigation: [The methodology of Justin and Gabriel's report](https://pulitzercenter.org/stories/suspicion-machines-methodology)
*For more insights, join the discussion [here](https://changelog.zulipchat.com/#narrow/stream/456003-practicalai).*
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:28Welcome to Practical AI. and database close to your users. No ops required. Learn more at fly.io.
0:42Welcome to another episode of Practical AI. This is Daniel Whitenack. I am the founder and CEO at Prediction Guard. We have some really exciting stuff to discuss today. So thankful to have my guests with us today, because there's a lot of talk about the dangers of AI or potential risk associated with AI, which we've talked about on the show. But I think maybe that kind of misses some of the actual real world problems that are happening with deployed machine learning systems that maybe have been going on for longer than some people might think. And maybe we can learn some things from those deployed machine learning systems that would help us create better and more trustworthy AI systems moving towards the future.
1:32So I'm really pleased to have with me today, Justin Braun, who is a data journalist at Lighthouse Reports, and Gabriel Geiger, who is an investigative journalist at Lighthouse. Thank you both for joining me so much. Thanks for having me. Thanks so much for having us. Yeah, yeah. Well, like I mentioned, I kind of teed us up to talk a little bit about maybe risks or kind of downsides of deployed machine learning systems. And you both have done amazing journalism related to what you've kind of titled here suspicion machines. and I think it would be worth, before we kind of jump into all of the details of that, which is just incredibly fascinating, if you could give us a little bit of context for both what you mean by suspicion machines and how this topic came across your desks and you started getting interested in it.
2:26Sure, I can start with that. I mean, the reason that we chose the suspicion machine as a title for our series is it's kind of a driving metaphor for what these specific machine learning models are doing within the welfare context. So a while ago, we wanted to investigate the deployment of machine learning in one specific area, but we're not sure which one yet. So in the US, there's been a lot of reporting about the use of machine learning or predictive risk assessments within the criminal justice system. Also in facial recognition and us over in Europe, looked at that reporting and noticed that there's a big lack of it over here in Europe.
3:03And so we were exploring different realms and settled on looking at welfare systems, which is a sort of quintessentially European issue, if you want to say. And in the last decade, welfare systems have become this sort of polarizing political battleground within Europe. How much welfare should we be giving out? Are people defrauding the state? How much money? And so we wanted to hone in on this one area to make it sort of manageable. And we decided to investigate the deployment of predictive risk assessments across European welfare systems. And basically what these systems do, I mean, they vary in sort of size and color, but the basic sort of mechanics remain the same, is that they assign a risk score between zero and one to individual welfare recipients and rank them by their alleged risk of committing welfare fraud.
3:52And the people with the highest scores are then flagged for investigations, which can be quite punitive and where their benefits can be stopped. So we landed on this metaphor of the suspicion machine because we felt that these systems were oftentimes essentially laundering or generating suspicion of different groups who were trying to receive welfare benefits that they needed to pay rent every month. And when you all started thinking about these suspicion machines, these deployed machine learning systems, were there existing examples, like concrete examples of how these suspicion machines were being punitive, maybe in either biased ways or just like in the kind of false positive error sort of way that's creating problems for people that it doesn't need to create?
4:43Was there actual evidence at the time, or was it just a big question because there wasn't any sort of quantitative measurement? So there was signs of it. So in the Netherlands, specifically, there was a case where 30 ,000 families were wrongly accused of welfare fraud. And it turned into this huge scandal called the childcare benefits scandal and eventually led to the fall of the government. And it turned out that the way that these parents were wrongly flagged for investigation was because of a machine learning model that the agency had deployed. But there was no sort of quantitative measure of what that model was actually doing.
5:18Nobody took it apart and actually looked inside and saw, okay, well, why was it making these decisions? Which is, you know, a huge reason why we as Lighthouse decided and were so adamant about the idea of we're not just going to investigate these systems or the classical journalist methods, you know, call people up sources, you know, getting contracts, but we actually wanted to take one of these systems apart. And that was sort of the big challenge or hurdle in our reporting. And I think Justin can maybe talk to what's sort of the existing literature on these predictive risk assessments. Yeah, I think my interest in the topic comes kind of from the broader discussions around AI fairness that really started after ProPublica published its machine bias piece six or seven years ago.
6:04And in the aftermath of that, there were a bunch of systems that worked in a similar way that were kind of discovered in various contexts. I myself worked a little bit on predictive grading systems. So during the COVID pandemic, some school systems replaced their previous, you know, handwritten exams with an algorithm that tried to predict based on previous exams, how well somebody would score in their final exam. And with each of these systems, the issue that emerges is essentially similar. Once you try to classify people according to risk and you have a training set that's not a perfect representation of the true population, you'll start running into issues like disparate impact for different groups, which is kind of the most hot button issue.
6:48But you'll also start running into how representative is your fairness data. In general, you'll start running into issues with, you know, where do you set the threshold? What values are you trading off when you set the threshold higher or lower? And so I was generally interested in that and then kind of joined Leidos at a point when Gabriel and some others had done a lot of the groundwork already to see whether there was something there in welfare risk assessments and then kind of took on the technical work from there. One question I have, just like even as a data scientist, just thinking like, okay, where do I start with this?
7:18This model is deployed by some entity. In theory, it's been developed by some group of technical either engineers or data scientists or whoever it is, where do you go about actually starting to find out, like, where does this model exist? Who has the serialized version of this model sitting on some disk somewhere in some cloud? Or like, yeah, where do you even start with something like that? So now there's a sort of trend of having algorithm registers where public agencies across Europe publish what different types of algorithms or models they're using. But that didn't exist when we started this reporting.
7:55So what we did was we made use of freedom of information laws in the US, I think they're called sunshine laws. And we started sending in these requests, trying to figure out at least where are you using predictive modeling within the sort of welfare system? Because, you know, you could be using it to look for fraud, but you could be using it for other things as well. And we started sort of slowly building out this picture of which countries were using predictive modeling at different places in their welfare system, and then sort of start slowly building a document base. So maybe we'd ask them for, we didn't start by asking a lot of times for like source code or final model files or training data.
8:33We'd start by asking for, can you give me like the manual for your data scientists for retraining the model every year? And that would allow us to ask for more specific documents and more specific questions like, okay, we know that there's a document called performance report 2023.html because we see it referenced in your manual for your data scientists. So we can request that. And then sort of built up to this place of, okay, now let's request the final model file, the source code to train it, ask for the training data, which we can get into because there's some prickly things there around data protection laws in Europe.
9:07So we kind of tried to do this tiered approach to sort of build for that final ask for asking for the model once we could make sure that our request was specific, because oftentimes agencies would try to resist our requests, saying they were too broad or we weren't being too specific enough, or trying to argue that disclosing certain documents could allow potential fraudsters to game the system. I've got to ask, as you did this sweeping look at how predictive analytics was actually deployed across Europe, even before we get into the specific case that you studied, are there any takeaways or trends that you saw in terms of how machine learning is actively being deployed by government entities or by welfare entities across Europe?
9:54Yeah, so I think it started essentially a bit later than in the United States. You kind of have this trend in policing, in kind of risk analysis. I would say that begins in the early 2000s, where you kind of have, I mean, semi-governmental organizations doing, credit risk scoring, the first kind of instances of predictive policing, also more serious thinking around big data mining for some risk analytics in the welfare context. And then I would say there's a bit of a bifurcation. So you kind of see some instances where big industry players, Accenture, the volunteers of the world, right? Like these big, companies hype up the case for big data analytics to be deployed across different sectors.
10:43And at the same time, you have a lot of failures when those tools are deployed. They often don't work very well. People who have to use them in the agencies don't know how to use them. You see some agencies that drop those systems. And at the same time, you see other agencies that kind of build up internal capacity and build those tools themselves, sometimes in collaboration with universities or smaller startups, but you kind of have these two pathways that continue to coexist at the same time. I would say in terms of the systems that we looked at, most of them were developed kind of from the early 2010s onwards.
11:19It's definitely gotten a lot more in the last five or six years. And across the eight or maybe nine countries now, I'm not quite sure how many we've looked at, but I think we've only seen a single country where we did not see evidence of predictive analytics being used to assess risk and welfare. Interesting. So I guess on the other side, I asked the question about evidence of these systems prior to your reporting evidence or cases where these systems maybe behaved in ways that caused harm or issues. On the other side, you mentioned this kind of hyped perception, potentially hyped perception of what these systems could do in a positive way.
12:00I mean, the main case for using these systems, as you mentioned, is to kind of catch fraudsters, from my understanding. On that side of things, is there evidence that, hey, yes, this type of fraud is a huge problem that we need to invest kind of advanced technology in solving? Or is that also kind of up in the air in terms of the, I guess I'm getting at the justification for using these types of systems on this sort of scale? This is one of the questions we try to address in our reporting a little bit. First of all, distinguishing between deliberate fraud and unintentional error is really messy and difficult.
12:38I mean, how do you prove intent? How do you prove that someone intentionally didn't report something? I mean, there's clear-cut cases where it's like criminal enterprises defrauding the welfare state using identity fraud. Okay, that's pretty cut and clear. But when it's individuals or family and they didn't report 200 euros, you know, is that intentional? Is that not intentional? How do you prove it? So that's already a challenge. What we did see is evidence of a lot of the larger consultancies tending to overhype the scale of welfare fraud and these estimations being criticized by, let's say, like academic studies.
13:08And, you know, when national auditors like the National Audit Office of France, for example, actually did, you know, random surveying to try to estimate the true scale of welfare fraud, they estimated at about 0.2 % of all benefits paid, whereas consultancies will estimate it at about 5 % to 6 % of all benefits paid. So there's a little bit of this situation where they're hyping up estimates to sort of sell the solution. At the same time, fraud does happen within the system and our reporting isn't meant to try to dispel the notion that fraud doesn't exist. But I think it's definitely still unsettled science on what the actual scale of welfare fraud is.
13:49and whether these systems that are being deployed in places like the case study we looked at are actually catching fraud or just catching people who have made unintentional mistakes and that these unintentional mistakes are being treated as fraud. To add on to that a little bit, I think the added justification that is often being used is that actually these systems are more fair than analog equivalents, that by using a machine you get rid of biases, and that they're better at detecting fraud than people are. And I think, as we'll probably get into later, there's good reasons to doubt both of those propositions.
14:24All of that was really good setup for this particular case study that I think you've highlighted in some of your recent work. I'm wondering if you could kind of set the context for the particular case study that you focused on, the particular model that you focused on, in light of what you were just talking about, about kind of scanning the environment, I guess, through information information requests and freedom of information requests to understand where things were deployed all the way down to like getting your hands on a model. So how did that transition happen? And tell us a little bit about the use case that you studied more deeply.
15:05As I mentioned earlier, we started by sending these freedom of information requests across Europe, eight or nine countries, and we started receiving a patchwork of responses back. So some places just said, no, we're not going to give you anything at all. Some places would be like, okay, we'll give you the manual. But then when you tried to ask for anything like technical, like code or a list of variables, they shot it down. But there was this kind of one exception in all of this, and that was the Dutch city of Rotterdam. And Rotterdam had deployed one of these predictive models to try to flag people as potential fraudsters and investigate them.
15:40And right off the at Rotterdam sent us the source code for the training process for their model. And we got really excited at first. We were like, wow, this is great. We started looking through the code and we noticed that when the scoring function in the code goes to load something called the final model.rds file, and we go looking through the directory and we notice, huh, wait a second, this final model.rds file, the actual model file that can be imported to the score isn't in the directory. So we emailed them back. We say, hey, guys, I think you made a mistake. There's this final model.rds file missing in the code directory, so we can't actually run anything.
16:20And they go, oh, well, yeah, psych, but you're not getting that one. And their justification for this was that if this was made public, potential fraudsters would be able to gain the system. So long story short, we went on this year-long battle with them to attempt to get this model file. and eventually the city, to their credit, decided to disclose this model file to us so we could actually run it. And what does this model do? I think Justin can do a good explanation of what this model actually does and how it works. Yeah, so it's a grading boosting machine model. It's a pretty standard machine learning model.
16:59It ingests 314 variables and it outputs a score. The issue that we ran into very quickly once we had access to this model is, well, what does this actually tell us, right? Okay, we can make up a bunch of people now and score them, but how do we then know what that means for those people? And so there were kind of two things that became important to figure out at that point. One was what do realistic people look like? And the second was what is the boundary at which a person is considered high risk. The second one was relatively easy to figure out. We kind of had some broad estimations of how many people are flagged each year.
17:42We could run some simulations and kind of see the distribution of risk scores. And at that point, we could take a good guess for what the threshold would be. Getting access to realistic testing data was a lot more challenging. And for a while, we thought we would have to just simulate a bunch of people, you know, take guesses. But But actually, Gabriel had requested some basic stats about the training data at an earlier stage. He essentially asked, look, can you tell us, give us a histogram for each of the variables so we can see what the broad distribution in the training data for ages, for instance, or for gender and so on.
18:19And our idea was to use those basic distributions to sample new people. But when I was meant to type all of this stuff down into a file so we could then run those simulations, I got lazy and I wanted to just scrape the document. And so it was an HTML file. So I opened it up and inspected it. And it turned out that the entire training data was contained in this file, which happens when you create plots with Plotly quite often. So if you want to leak something to a journalist, that's a good way to do it. There you go. Yeah. So we got a backstone and got access to the entire training data. At that point, the question became, okay, now we know what realistic people look like.
18:57What tests can we actually run in terms of figuring out who does this model flag at higher rates? Does it have justification to do so and so on? And the one thing that was missing from the train data was that we didn't have access to the labels itself. So we knew, you know, your age, your family background, your job history, that kind of stuff. But we did not know if you had actually committed fraud or not. And that meant that, and this is the big limitation of our story, but that meant that we could essentially only understand which characteristics lead to higher or lower scores, but we wouldn't know if those scores are erroneous at higher rates for one group rather than another.
19:33So I just want to be very open about that. That is a limitation of the design. But having access to the train data, having access to the source code, being able to see how the train data is constructed, having access to the final model file, all of that allowed us to investigate a bunch of aspects with the system, which I think still made for a very valuable story, both in terms of explaining how this stuff works, but then And also in terms of showing that there are likely consequences which seem to be discriminatory against certain groups. Probably a lot of our listeners will be familiar with what a gradient boosting machine is.
20:04And sort of like maybe this is one of the tutorials that you ran on a Jupyter notebook when you were first taking your kind of data science 101. So the model, I think, is very familiar. I think a lot of the interesting things here are related probably to the model features and that sort of thing. Did anything jump out to you maybe even before you kind of ran a kind of larger scale analysis in terms of like the features that were included in the data set and how those may or may not like intuitively be connected to this sort of welfare fraud situation? Did anything jump out when you were kind of doing your initial discovery and kind of exploratory data analysis on this data?
20:51Yeah, for sure. Though I think it's maybe important to preface this with saying that including features that seem discriminatory does not automatically lead to discriminatory outcomes. And I think that is sometimes being confused, right? You can get discriminatory outcomes without features that look bad, like, I don't know, racial or gender, racial background or gender or something like that. But it also works the other way. You can include a bunch of these features and not get any discriminatory outcomes, right? Both of these things are possible. That being said, there were a bunch of features that seemed perfectly reasonable.
21:25You know, contact with the welfare agency. How often have you been there? Have you missed any of your appointments? That kind of stuff. There were a lot of demographic features. And I think those get into trickier territory. Some like age are maybe justifiable on some level. Gender gets a bit harder. And then a lot of features measuring through proxy, but measuring ethnic background through language skills. I think there was 10 or 12. Gabriel, correct me if I'm wrong, but definitely a lot of variables on language skills. No, I think 30 or something. Oh, 30. Yeah. Yeah. Because it measured everything from like your Dutch, like spoken Dutch fluency, writing Dutch fluency, the actual language you spoke.
22:03So there was like a categorical variable with like 200 values or something. So it got as granular as the specific language you spoke, whether you speak more than one language. But anyways, continue, Justin. Yeah. And then I think in some way, the weirdest set of variables were essentially behavioral assessments by the caseworkers. So we actually got access to some of the variable codebooks. And in there, it said that there was a variable that essentially where people were meant to judge how somebody was wearing makeup, especially for women. So stuff that just seems really sexist. So those variables were included, which is problematic in and of itself.
22:42But then the way they were transformed in the preprocessing steps was that essentially this textual data was just transformed into a zero one variable, depending on whether there was anything in this field or not, which is also, I mean, you just lose a bunch of maybe the more interesting information if you do that. But I think that set of variables, because it's just based on individual caseworker assessments, if your claim is that the system should lead to reduction in bias, and then you include these variables that are so obviously subjective, I think that kind of undermines your claim right away.
23:15And in the data set, like in terms of the label and the output, were you able to understand at all, like, oh, these are investigations that happened that actually were verified to be fraud or not, essentially like a one or zero type of label? Or how is that set up? Yeah, so we did not have access to the label, which is, again, the big drawback. So we could only score people who we know they had labels, but we didn't have that label ourselves. Gotcha. But we did a bunch of ground reporting to essentially work around that. And maybe Gabriel can speak a bit to that. Yeah, I mean, two things first.
23:52I mean, just to talk about how the training gate is constructed first. It's over 12 ,000 past investigations that the city has carried out. And these past investigations are not a random sample. So there's some subset within there that's random, I think about 1 ,000. But all the rest of the cases are just where investigators have looked at the past, either through anonymous tips or through these kind of theme studies that they do where they say, this year, we're going to check every man living in this neighborhood. So it's not a random subset of people that they're training this model on, which is problematic in the first place.
24:23The second thing is that this label, yes, fraud, no fraud, doesn't distinguish between intentional fraud and unintentional mistakes, right? So these are flattened into the same thing when labeling the training data set. So those are, I think, two problematic things right off the bat. I think an even third, more complicated thing is that the law for what is considered fraud has actually changed over time. And this training data spans back 10 years. But all that aside, one of the things that we wanted to do with this reporting was to look at the impact of being flagged for investigation. What does that mean for a person?
25:01And how are they treated by the system? And so we did a bunch of ground reporting in Rotterdam, and we sort of used the results from our experiment to build profiles who would be considered some of the most high-risk people. And we saw that it was, you know, one of them at least would be like single mothers of a migration background who don't have a lot of money, financially struggling, living in certain majority ethnic neighborhoods. So we did a bunch of ground reporting in those places and found people, and it was quite challenging. People were quite afraid to talk. People who had been investigated in the time span that the model was active.
25:36And what we found was that they were treated incredibly punitively by these investigations from the city where fraud controllers are empowered to raid your house at 5 a.m. in the morning unannounced, count your toothbrushes, sift through your laundry, go through all your bank statements, and that even the smallest mistakes, like forgetting to report 100 euros, could leave you landed as an alleged fraudster. So I think there's even, you know, based on reporting, there's reasons to even question the validity of the label and the consistency of the label. But beyond that, I think what we established for the reporting is that the consequences of being flagged, even if in the end, you're found to be completely innocent, just having people, you know, raiding your house at 5am, asking you questions about your romantic life in front of your children.
Read the full transcript
26:26I mean, that's a negative consequence in of it itself, even if you're found to have done nothing wrong.
26:55This is a changelog news break. The biggest product news out of OpenAI recently is GPTs, custom versions of ChatGPT that you can create and sell for specific purposes. You build these GPTs by crafting special prompts that are fed to ChatGPT prior to it interacting with a user. Is it any surprise that crafty technologists have convinced ChatGPT to spit out a bunch of these custom prompts via prompt injection? I wasn't surprised, but I was a bit delighted to read through the collection of GPT prompts to see what they're made of. This Gen Z 4 meme prompt, which helps you understand the lingo and latest memes that Gen Z are into, is kind of hilarious.
27:40Quote, speak like a Gen Z. The answer must be an informal tone. Use slang, abbreviations, and anything that can make the message sound hip. Especially use Gen Z slang as opposed to millennials. The list below has a list of Gen Z slang. Also, speak in low caps. End quote. Low caps, more like no cap. Am I right? I'm so old. Fair warning, though, from the collector of these leaked prompts who says, quote, there is no guarantee that these prompts are the original prompts and these leaked prompts are for reference only. You just heard one of our five top stories from Monday's Changelog News. Subscribe to the podcast to get all of the week's top stories and pop your email address in at changelog.com slash news to also receive our free companion email with even more developer news worth your attention.
28:30Once again, that's changelog.com slash news.
28:50All of this is very interesting to me from a data science perspective, because a lot of these things are kind of, yeah, things that I know we've talked about, you know, on this podcast, but also in my day-to-day work, things that have come up that you sort of establish as, you know, best practices around how you construct your label, how you construct your features, like in a responsible way to do well at your data science problem. I do want to get to the actual like model performance here in a second, which is one question is like, well, we see all of these flaws in the data. Does the model actually work or have all of those kind of underlying problems poison the output?
29:30But I think before then, I'm just wondering, like as a person who provides occasionally consulting services to other people in data science, did you get a sense at all for like the city of Rotterdam hired X consultancy to give them the model that they deployed and are using is just sort of like the consultancy through the model over the fence and like, here, use this? Or how much interaction was there with actual Rotterdam employees? And how deep was the understanding of how this model was built and deployed? Or was it just a sort of contract? Here's money. Here's the model. All right, let's put it into production.
30:10What was the interaction like there? Were you able to discern any of that? Not super deeply, but from what we do know, the city put out a tender asking for someone to come in and build a predictive model for this purpose. Accenture won that tender, put someone on it, and there was a Rotterdam data scientist involved, but who presumably, or from what I can tell, didn't have any sort of machine learning background. I'm just a normal famous scientist at the city. Rotterdam set up the whole code base, trained the model, developed all the code for the pre-processing, trained the model, handed it over to the city, and kind of went by, like, we're gone now.
30:49And from that point on, Rotterdam took full control of the model. Like, they would retrain it every year. They made, like, adjustments to, like, for features, and also decided to exclude some features like nationality. But I do think that during that time, Rotterdam upgraded its own data science capacity. So by the time we got there, they did have like two people who were specialized in machine learning that were looking over the model. That's my understanding of the basic setup. Yeah, super interesting. I do want to get to the kind of model performance, I guess, because I know this is something that I've got asked when I've done workshops and I talk about either like fairness or bias in models.
31:30There's always someone that kind of comes up with the question of like, well, if the data is biased, but I'm still like the model's accurate and I'm predicting accurate results, is that a problem? I think there's problematic things about that, how you might answer that question in and of itself. But in your case, was the model actually helping in any way? Or were the problems kind of so deep in the data and the way that the labels were generated such that the majority of what it was producing was maybe more chaos or issues? So in the test set that the city used, and we have kind of their documentation of that, even though we don't have the labels ourselves, we see that in the set, there is a 21 % baseline rate of fraud or some kind of wrongdoing.
32:19And the model, kind of depending on where you set the threshold, but the model essentially has a hit rate of 30%. So out of the people selected, around 30 % of them are labeled within the positive class. So it's a 10 % improvement above random. Is that good? Is that bad? The ROC curve looks absolutely terrible. Margaret Mitchell, who many listeners probably know, called it essentially random guessing. I'm not quite sure if I would go that far, but it's certainly not anything to write home about. And we see that there's huge disparities in who's getting flagged in their characteristics. Does the label data show that there's a reason for that?
32:59Maybe. But because we have some idea about how the training data was constructed, specifically through these theme investigations, there's a very strong probability that a lot of these patterns that we see in terms of who's getting flagged is a function of the selection process that leads to somebody being included in the training data rather than of actual fraud being committed. I can give an example of how that might work. Most of the men in the training data very likely were selected through one of these investigations where all men in a certain neighborhood were investigated, which have a pretty low likelihood of actually finding fraud.
33:32That kind of implies that most women were selected by anonymous tips or random sampling. and those things have somewhat higher probabilities of detecting fraud. And so if your method of selection impacts how likely it is that the person who you investigate has actually done something wrong, then the training set that you train your model on will contain patterns that are a function of your selection method rather than of the real world and how fraud patterns look in the real world. And so we couldn't conclusively prove this because we didn't have access to who was labeled or who was selected, like within the training set, we couldn't say who came from which source, but we know that these different sources fed into the training set.
34:13And it seems very probable that this type of selection method would lead to these kinds of disparate outcomes. I think there's all sorts of things to learn in this story as even just a data scientist setting up data sets and trying to train models. You know, I come, of course, from a certain perspective in kind of what touches me about this story. And I'm so glad that it's out there and there's some transparency around this. I'm wondering, could you speak a little bit to the reception of this story, maybe more widely by non-technical audiences in terms of realizations that people were coming to or responses that came out of people realizing how these systems were constructed and how they perform in reality versus maybe what their perception was prior.
35:08Kind of to answer to that question, I think, first of all, one of the big goals of this project and the piece that we published with Wired, where we kind of take leaders through the model, how it works, was to have it be an educational piece of journalism too. Like you've been hearing about machine learning and the sort of impact it has on your lives, but very few stories actually take you through like the full life cycle of our model. What does it look like, quote unquote, inside the machine? So we really wanted to make an educational piece in that sort and also talk about, you know, what Justin has covered.
35:37What are the different sorts of problems or flaws in the system? What are the consequences of those flaws? And, you know, I think normal people, of course, found the sort of discriminatory angle or the fact that, for example, like single mothers were penalized more or, you know, I think that was something that they took away from. But surprisingly, one area that surprised me a little bit that people seem quite fixated or curious by was the decision trees portion. So what we try to do in that portion of the piece for people who haven't read it yet is we take some decision trees from the model, from this gradient boosting model.
36:12We show how this creates nonlinear interactions, right? So features have sort of relation, in fact, each other differently relationally. So, you know, in decision tree X, if you're a man, you might go down the right side of the tree. And if you're a woman, you might go down the left side and you will be evaluated by different characteristics. So that was seemed to be something that really like seemed to resonate with leaders, like questioning, like, OK, well, that's this is how it works. And, you know, is that fair to me or, you know, it makes it difficult for me to understand how these interactions work on a political level.
36:43You know, Rotterdam, to their credit, was quite graceful when we presented them with the results. And they sent back the statement saying essentially, they called our results informative, educational, which in the field of investigative journalism never happens. Like someone say, the subject of your investigation saying it's informative and educational is I think never happened. When they're the subject. Yeah. And called on other cities to do what they have done, to be transparent. And I found that incredibly brave and elegant response to what we've done. And they were sort of debating whether to continue the use of this model and then decided that they weren't going to use it anymore, that the sort of ethical risks were too high.
37:32And then I think, I mean, elsewhere, I don't know, Justin, if you have any reactions that stuck out to you. Yeah, maybe the one thing that I would add is that I think this field of algorithmic accountability reporting, but even the academic discussions around it has, I don't want to say suffered, but it has been kind of constrained a little bit by a streetlight effect following machine bias, right? You had this big story coming out. And then afterwards, for years, everybody was talking about these various outcome fairness definitions. And I think that's a very valuable debate. I myself almost enjoy it.
38:02Like I think it's some of it is just mathematically very interesting. It's really difficult ethical questions that it brings up. But I think a bunch of the other dimensions of fairness in the life cycle of the system have been neglected. And Gabriel and I, in the past year, have kind of been making the rounds and the cases to people that we should be looking at algorithmic fairness more holistically. We should look at the training data, we should look at the input features, we should look at the type of model that is being used and how that maps onto our understanding of the process. And then we should also look, of course, at the outcome fairness stuff.
38:34But I actually think, and your reaction kind of spoke to that, I think this training data bit is is probably the most interesting one. And one that I have both academic training as a computer scientist and also as a political scientist. And when I took my computer science classes, nobody ever talked about how do you set up a representative sample? I was kind of like, we take whatever data we have and then we try to run as many models over it. Use it all and all features. Right, right. And well, that might kind of up your performance along certain metrics, right? On some level, if the data doesn't contain the functional relationship that you're trying to model, you can't get there.
39:10And I think that's a lesson that I hope some of the, yeah, maybe practitioners who read our piece also take away from it. Yeah, that's super helpful. I think you got to where I wanted to ask anyway, because I know we have listeners that are practitioners and are probably thinking to themselves, like, what is a kind of takeaway that I can take away from this? Because I would say from my experience, at least most data practitioners are not intentionally trying to create harmful outcomes from their systems. They do actually want to be responsible. It's just sometimes they might be somewhat confused or constrained in certain ways that don't allow them to spend time thinking about those things.
39:54But yeah, I really appreciate you bringing us around to that as we kind of close out here and we look maybe to the future. We started out this conversation, I kind of mentioned, you know, there's all of this talk, of course, constantly swirling around us about the dangers of AI and all of that stuff, which is operating on multiple levels, some of which are useful and some of which aren't probably. But I want to ask both of you, maybe as you look towards the future, post this project, what you've done here, what's on your mind as you look towards the future of how this technology is ever expanding?
40:30What gives you pause? What gives you hope? What do you hope people are thinking about as we kind of look to the future in how this technology is developing? So there's two things I would respond to that. One is that I hope we'll have more discussions around transparency around these systems. I think that's a precondition for anything else. And for that to happen, there's an argument that needs to be dispelled. And that argument is that making these systems public allows people to game them. One, I think it's really, really hard. And there's some very good academic research that shows how hard it would be.
41:06And two, well, these systems operate essentially like bylaws, right? They're essentially administrative guidelines encoded in a model file for how a decision is being made in some bureaucracy. And I think it's really hard to make the case that such guidelines should be secret. And so, yeah, I think we need to have a discussion and make the case proactively that transparency in this space and encouraging people to learn how they work is a good thing. And encouraging people to game those systems is probably a good thing because that means you're probably closer to abiding by the law. And if you can game the systems, then maybe they aren't very good.
41:43That's the first thing I want to say. The second one is that most of the systems we've looked at are pretty terrible in most ways. I They think they don't work very well. They have either, you know, use features that are absolutely terrible or have training data construction that is really problematic or, you know, have disparate impacts on various groups. Almost every single system we've looked at so far has one or multiple of these features. But there are some systems that maybe are better. And it's possible, I think, if you think very seriously about how you do each of these steps, the feature selection, the training data, and then constructing the model and then evaluate for bias and then potentially retrain, reweigh your training data and so on.
42:21Maybe it's possible to get to a better place. Technically, it certainly is. And I think then you get to a different set of questions. And I hope that the conversation at some point can move beyond kind of the gross incompetence in a way which we are showcasing across the board, but can move to a place where we can discuss, okay, let's take this best case scenario. We have a system that doesn't have obvious bias and so on that was constructed carefully. Should we do this? is it a good idea? Is a machine making the decision removing something inherently kind of valuable from this type of interaction?
42:55Is the machine actually more explainable than a human is? And is that a good thing? Is it equal treatment because everybody's being scored by the exact same system and not by individual caseworkers? Or is it not equal treatment because the tool contains a decision tree-based model, and so different people are based on different characteristics? How do we think about systems that include some level of probabilistic assessment? it? Is that something that we think an administrative decision should do? And then, of course, we can also have the maybe fun for some people discussions around like which fairness definition is the best, whether we should seek to minimize or equalize false positive rates across different groups and so on.
43:33I think there's a bunch of really important questions that society has to grapple with here, but I don't think we're there quite yet in most cases. And so long as we aren't, I think Gabriel and I will have plenty of work showcasing incompetence and all that stuff. But I hope that at some point we can move beyond that. Yeah. Anything to add, Gabriel? No, I think Justin summed it up really well. I'll just kind of tease that we do have some reporting that's coming up in the coming year that will, I think, grapple with some of these thornier ethical issues. You know, ask questions like when and if ever, is it okay to use these systems?
44:08I think maybe one thing that I will add, though, is I think it is important for people like practitioners that are listening to your audience to also take a step back and to maybe not see always the deployment of these systems or these sort of thorny fairness questions as like a math problem, but it can also be a sort of wider societal problem as well. So for example, in the European welfare context, we've seen in everywhere we're looking models that attempt to detect fraud. But what we don't see is models that try to find people who are eligible for welfare benefits who aren't using them because they're afraid of the system.
44:45And we know this is a huge problem. In places like France, 30 % of people eligible for welfare don't use it because they're scared of the system. This has consequences for people not using welfare, but also has consequences downstream for society. So imagine families that aren't able to feed their kids, developmental issues that come from that. So I think it's always important and something we try to raise in our reporting to kind of take a step back and ask, you know, should we be doing this? to think about the premise of why are we actually deploying this model and to rethink that. And at some points and think about, you know, is there a better way to use this technology or are we only kind of narrowing in on one piece of this picture?
45:24Yeah, that's great. I think that's a really wonderful encouragement to end things with. We will certainly be on the edge of our seats looking for your future work. And I encourage everyone, we'll include the links to Gabriel and Justin's work in our show notes. So I encourage you, go and explore it. There's lots of great graphs and references and even more technical description of the methodology than we had time to go into here. So dig in and learn about what they're doing. It's really wonderful. And yeah, thank you for your work, Justin and Gabriel. And thank you for taking time to join us. Thanks so much.
45:59Thanks so much for having us.
46:09Thank you for listening to Practical AI. Your next step is to subscribe now, if you haven't already. And if you're a longtime listener of the show, help us reach more people by sharing Practical AI with your friends and colleagues. Thanks once again to Fastly and Fly for partnering with us to bring you all Change Talk podcasts. check out what they're up to at fastly.com and fly.io and to our beat freaking residents breakmaster cylinder for continuously cranking out the best beats in the biz that's all for now we'll talk to you again next time
From the publisher
In this enlightening episode, we delve deeper than the usual buzz surrounding AI’s perils, focusing instead on the tangible problems emerging from the use of machine learning algorithms across Europe. We explore “suspicion machines” — systems that assign scores to welfare program participants, estimating their likelihood of committing fraud. Join us as Justin and Gabriel share insights from their thorough investigation, which involved gaining access to one of these models and meticulously analyzing its behavior.
Changelog++ members save 3 minutes on this episode because they made the ads disappear. Join today!
Sponsors:
- Changelog News – A podcast+newsletter combo that’s brief, entertaining & always on-point. Subscribe today.
- Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com
- Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs.
Featuring:
Show Notes:
Something missing or broken? PRs welcome!




