#408 First Look: Open AI cuts diagnostic errors by 16%. Or does it?

30 Jul 2025 · 52 min · 17 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

A “first look” at an OpenAI preprint with Penda Health claiming diagnostic errors fell 16% (and treatment errors 13%) across ~40,000 patient visits in Kenya, and a critique of how such medical AI claims were publicized before peer review.

Guests

Dr. Paul Wicks, neuropsychologist and journal editor; long-time digital health evidence critic.

Key claims

The episode argues the preprint’s evidence foundations are weak: unclear adherence to AI/clinical quality reporting guidelines, possible post-hoc/statistical “story shaping,” and limited inter-rater agreement among ~120 independent physician reviewers (only ~1/3 from Africa). It also highlights that the “16%” metric is based on subjective chart/documentation judgments and that LLM components may add variability. The host notes a pattern: Microsoft and OpenAI both promoted preprints via media/LinkedIn, despite norms against press releases.

Notable examples

comparison to Microsoft’s NEJM-vignette claims; analogy to antibiotic stewardship/statin alerts when OpenAI compares effect sizes; mention of a parallel Gates-funded RCT (diabetes/hypertension) not emphasized in the preprint.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Analyzing OpenAI's Claims

0:45 to 3:14

Discussion about OpenAI's claim of a 16% reduction in diagnostic errors.

“The most good one can do as a doctor in a lifetime, you know, there's estimates that, you know, you might save half a dozen lives or something with the decisions you make.”

Concerns About AI in Healthcare

3:14 to 6:06

Exploring the implications of AI in healthcare and the risks involved.

“They are a way of spreading your methods.”

Understanding Preprints and Their Risks

6:06 to 10:33

Discussion on the role of preprints and the responsibilities of researchers and media.

“Or over these people who publish their methods, you know, they said they had some missing data, but they're accounting for it.”

Evaluating OpenAI's Methodology

10:33 to 14:01

Critical assessment of the methodologies used in OpenAI's claims and their validity.

Analyzing Diagnostic Errors Reduction Claims

14:01 to 19:16

The discussion centers on the statistical analyses of AI's role in reducing diagnostic errors and the challenges in interpreting these results.

“Or perhaps they sort of, they ran it and they ran it again and they ran it again and they ran it again.”

Challenges in Evidence and Bias in AI Studies

19:16 to 24:25

An exploration of the challenges in AI studies and the biases that can arise from them, especially regarding claims of improvement.

“which was an interesting methodological application.”

Quality Improvement vs Randomized Control Trials

24:25 to 28:04

A contrast between quality improvement studies and randomized control trials, discussing ethical considerations and methodologies in AI healthcare applications.

“Because we're seeing big announcements from the US AI strategy, right?”

Quality Improvement vs. RCTs in Healthtech

28:04 to 29:18

Exploring the differences between quality improvement studies and randomized control trials in healthtech, emphasizing the importance of rigorous protocols.

Concerns on Research Transparency

29:18 to 30:21

Discussing the lack of transparency in the funding and authorship of studies, and its implications for credibility.

“So really well-designed, elegant RCT that's not mentioned at all in OpenAI's preprint, even though it shares at least one author in common.”

Speculation on Influencers in Health Studies

30:21 to 31:07

Speculating on the potential influence of Bill Gates in health technology studies, raising ethical concerns.

“And so when I saw this Gates-funded RCT, I expected to read, oh, we also have this nice grant from the Gates Foundation.”
Show all 17 chapters

Ethical Oversight in AI and Health

31:07 to 33:15

Analyzing the ethical oversight in AI applications within healthcare and the implications of unequal power dynamics.

“It would seem like a good time to go, look, great.”

Algorithmic Risks in Diverse Populations

33:15 to 35:34

Discussing the risks of algorithmic bias in healthcare AI, particularly concerning diverse populations.

“And I can tell you with absolute certainty, if you were a pharmaceutical company and you did that, you would be in heaps of trouble from large regulators with significant hours to compel you not to do that again.”

Trust and Transparency in AI Healthcare

35:34 to 41:25

Emphasizing the need for trust and transparency in AI systems used in healthcare, discussing the pitfalls of premature claims.

“They have$50 million programs of research with Stanford and Harvard and all the big places.”

Trust and Evidence in AI Healthcare

42:04 to 43:37

The importance of trust and evidence generation in AI healthcare tools is discussed.

Advice for AI Tool Development

43:37 to 47:45

Guidance on generating evidence for AI tools in healthcare and the need for collaboration.

“And then there will be other local groups in the US and other places.”

Ethics and Patient Safety Concerns

47:45 to 50:28

The ethical implications of publishing AI research without proper validation are examined.

The Role of Big Tech in Healthcare

50:28 to 52:16

A discussion on the influence of big tech companies on healthcare evidence and practices.

“yeah so i think um look i think i i predict that there were voices within each of the company that said this.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Dr Paul Wicks:Hey everybody, welcome to another episode of the Health Tech Podcast. This is a first look and this is a first look at OpenAI's new pre-print with Penda Health saying that diagnostic errors, and I'm reading this, have been cut by 16 % across nearly 40 ,000 patient visits in Kenya, which obviously sounds huge and amazing and wonderful. And to discuss this based on a very impressive and very in-depth and great, frankly, LinkedIn article. I've got Dr. Paul Wicks here, neuropsychologist, journal editor, long-time sleuth of digital health, I will say, protecting us from all things research and evidence.

0:48Dr Paul Wicks:And he's published this article, extraordinary claims weak foundations is essentially the vibe um that's a bit of a teaser but paul welcome to eltip podcast how you doing mate thank you i'm really good uh it's great to be here and glad to be talking about this uh this fascinating preprint yeah so open ai a big beast um lots of money doing lots of things uh arguably the most important or relevant or something company at the moment um redefining the world that we live in they have a huge amount of power and they've published something and so everyone's going to stand up and listen everyone's going to sit and watch or read or take heed and so a classic spider-man great power great responsibility etc um why is this important and why did you jump on it so i think this is important For the reason you say, there's the scale here of these large tech companies, these large AI companies.

1:55The most good one can do as a doctor in a lifetime, you know, there's estimates that, you know, you might save half a dozen lives or something with the decisions you make. The most harm you could do, and not to be, you know, good, but things like cases like Harold Shipman, cases like Lucy Lett be terrible, but sort of show you the sort of maximal damage that sort of one bad actor in a system can have. however the scale then cuts both ways when we come to health applications imagine the harm of an algorithm that we invisibly trust imagine the harm of an algorithm where as a result of it being ubiquitous in health deployments around the world maybe we take big decisions maybe we stop training so many doctors maybe we start de-skilling other healthcare professionals on the basis that the algorithm's got it under control we're at a really important time and if we think of our sort of future colleagues 20 years from now, some of the foundational decisions we make, some of the foundational processes we build to safeguard that system are going to be the bedrock of whether or not their patients are safe, if their patients are well cared for.

2:57Dr Paul Wicks:So what have OpenAI said that they are capable of? Yeah, so I want to sort of cite this in a context of something that had just happened two weeks earlier as well. So as you know, there's a relationship between Microsoft and OpenAI, and Microsoft has made a large investment into OpenAI. And just two weeks previously, Microsoft had put out blog posts, headlines, and many of their senior staff on LinkedIn were making specific claims that their technology, as a result of running through some vignette cases from the New England Journal of Medicine, were able to make some really difficult diagnoses.

3:39and some of the headline figures there was that you know boiled down their technology was 80 % accurate compared to some human doctors who are only 20 % accurate and in that case what we saw was a preprint but the way preprints are meant to be used that they're meant to be used by scientists to get out something that is so super exciting that you can't keep it to yourself and you want to share it with the community you want to show your methods and you want to get feedback from the community. They are a way of spreading your methods. They were extremely useful in COVID when we didn't have time to wait for, say, long peer review cycles from typical journals.

4:19But there is really only one rule. And the rule is don't broadcast it in the media when it's a preprint.

4:26Dr Paul Wicks:Interesting. And Microsoft did that, and they went and gave interviews. There was press about it. And so I first came to the Microsoft paper from the perspective of an associate journal editor thinking oh gosh this huge corporation has just violated this principle of preprints if i got this across my desk could i even send it out for peer review and and i started with that angle and then actually as i started reading it um i started putting that apart so so i first started with this linkedin article about microsoft ai and then just two weeks later open ai did the exact same thing and i'm not suggesting that you know microsoft and open ai need to sit awake all night wondering what i'm thinking about in the middle of rural staffordshire um but like it's starting to feel like a pattern right where you put out a pre-print which is not meant to be press release and not meant to be marketed and these two companies which all together make up you know a significant chunk of the stock market and and you know us ai power um have both done the same thing in the same area which is medicine um now the open ai paper is from a different part so the new england journal of medicine paper is entirely theoretical right they've got theoretical vignettes no patients touched by the system the open ai systems deployed down in kenya uh down in nairobi with a a private um social um social enterprise health system and and it's been live it's been deployed on like you said earlier tens of thousands of individuals but again the pattern goes um along the lines of well we've seen this big press push but then uh when you start reading into the details a whole bunch of questions emerge that sort of looked like maybe this pre-print is not quite as fully baked as we might expect and and in both cases both studies were missing some really fundamental underpinnings that make research good that make it replicable that make it safe that make it reliable and allow us to trust right that allows to trust different organizations and trust each other as a researcher So where this all comes from is, you know, we as researchers or clinicians or whatever we are, we always default, I think, to giving each other the benefit of the doubt.

6:37Yes. You know, Dr. Rector's made this referral. I'm sure they're trying their best. Right. That comes in. Or over these people who publish their methods, you know, they said they had some missing data, but they're accounting for it. I'm sure they're trying their best. That benefit of the doubt is not an automatic right. and i think when you come from a background of academia or hospital research whatever you some some of that kind of comes along with the territory and what this made me question was when you have a big tech company that is worth hundreds of billions of dollars that is moving whole economies around that's building data centers that is part of foreign ai policy

7:14Dr Paul Wicks:do they get that same benefit of the doubt or not and are they taking actions now that earn that benefit of the doubt so that as they build up to this mass scale we understand the process we understand what's gone into it because there has to be an appreciation of intent right because that benefit of the doubt is closely related to trust and intent and i think when you've got i mean it's a company is faceless nameless and of course obviously we know who sal maltman is but an entity like that you can't you can't really appreciate intent as much it's far more difficult to grasp I guess and so before we move on to like what they've actually done here when you mention pre-prints and I think this is a really really really good learning point for people that are in health tech digital health or even students of health tech and digital health perhaps clinicians even whose rule is that that you don't put preprints in media because this is a really interesting section of this for me and where we sit in the ecosystem at at the border of health tech and media is it you mentioned rule you mentioned the word principle as well and i'm wondering whose responsibility is it is it the responsibility of the people that have done the thing to not put it into media or is it responsibility of media and a journalist in the source of the search for truth to not publish the pre-print or is it both that's a really great question so so some of my context for this uh comes from when i i sat on the board of the BMJ that was involved in setting up something called MedArchive.

9:10So preprints have been around for a long time in biology, in computer science, in physics. And they go on the basis that if you've done something interesting in the Large Hadron Collider that is going to change the field, you should tell everybody now, right? But academic peer review journals take at best six months from when you submit to getting accepted. The desk rejects a lot of stuff. The more famous the journal the more stuff they reject so the bmj rejected about 97 of what was sent to it so if you have a great medical breakthrough and you send it to the lancer the bmj nature etc and every one of them's taking a few months to reject you you know you can talk about it a two-year gap in some cases so so so when there is a clinical decision to be made an example during covid would be if hydroxychloroquine doesn't work and we have a good evidence base from recovery to do that it's more useful in a pre-print we can take this immediately out to the wards maybe an unusual situation you know we know it really takes 17 years for most science to get from um you know bench to bedside but this was an unusual case so at the bmj at harlan krumholz from yale who was the architect of med archive was making the case that that you know it's it's important that we move with speed um but that these preprints will eventually be peer-reviewed and the benefit of putting them in a good archive is they get something called a digital object identifier and that allows you to track the version number all the way through so if and actually so there's a big theme here of basically like say what you're going to do do it run the analysis and then let other people look at it and so the point of getting a pre-print out is you could get early comments on it you know and if you've done something really bad like plagiarized a figure or done something dangerous or what have you that could be detected even before it went into a peer review yeah and the community that does peer reviews in their spare time um might add something useful and maybe you update it to a version two or version three before you submit it to the journal so um but the big issue that people were concerned about was that people could or academic researchers or clinical researchers could put something dangerous into a preprint something false you know hydroxychloroquine cures covid off you go and uh so there were lots of disclaimers made one of them was a big disclaimer saying do not use this for clinical decision making but another one was don't press release stuff and so you know the journals themselves at first were hesitant about these because it potentially disrupts their entire business model and so these are the groups that have said guidelines uh so cope as the overall global body about publishing ethics has guidance on it but it's up to journals like nature etc to say well look you shouldn't be doing marketing of a preprint now what are the

11:56Dr Paul Wicks:consequences what is the sort of case law nobody knows but it doesn't sit right because how many times have you peer-reviewed a paper where you go oh there's something wrong in the methods and the authors go oh yes sorry that was wrong we'll update it so in this case if the claim is that it's 16 better okay what if that doesn't survive peer review what if it doesn't hold up what if it's 14 what if it's 40 what if it's minus two right like that's that's kind of the sorts of questions that it started kicking off so let's so let's go into that then so having set the the scene here that this is a pre-print and principally that is not okay to be producing this it has been produced and it's been widely reported on what it what have they what have they reduced by 16 exactly what are they claiming to have reduced by 16 it's a little hard to tell and in part it's a little hard to tell because in the ideal scenario you publish quite a thorough protocol that says what you're going to do before you report the results and protocols should follow guidelines and the guidelines are there to ensure that we compare apples to apples and oranges to oranges.

13:16So if you take Scribe software, which is very hot at the moment, there are guidelines out there about what is a good way of measuring quality or how do we report that so that what you avoid is company A reporting stats that make their tool look good. Maybe it's very quick. Company B making their tool look good. Maybe it's very cheap, et cetera. We need to compare that with like. So there are these guidelines that either take existing clinical trial guidelines or extend them for the special additional work that needs to go into place with AI to help other people understand what was done. And the work that OpenAI describes with Penda Health doesn't have that.

13:53So it just sort of shows up and says, this is what we did. And the challenge with that is that we don't know if the statistical analyses we're seeing were planned all along. Or perhaps they sort of, they ran it and they ran it again and they ran it again and they ran it again. And eventually it sort of looked like the story that you wanted to tell. and now you tell the story, right? So that's what those editors say. Now, there's no bad intent in that. All researchers fall in love with those ideas. We hope the software does better. You know, we always want it to do better, but we have to be open to the idea that it might not work out that way.

14:30And so that's why these are so important. And so I was surprised in both the Microsoft AnyGM AI paper and the OpenAI Penda Health preprint that there either was no guideline adhered to, or in the case of the Penda paper, they said they adhere to quality improvement guidelines. But when I read them, when I read the guidelines and I read the paper, they didn't match up. So, you know, talk about that benefit of the doubt. If you sort of start off with one exclamation mark, you know, the exclamation marks are sort of making a cue outside my head as I go end to end through this 160 odd page paper. And I ended up feeling not reassured.

15:10Dr Paul Wicks:And so they talked about, So it's a 16 % relative reduction in diagnostic errors and 13 % in treatment errors. It's, I don't know, again, I work in media, right? And there's a load of stuff that you have to do in order to make something sound even vaguely interesting to news. You look at this and you go, well, 16 % is not a huge number. So you go, okay, well, how do we make this newsworthy? we need to have some punchy words here some punchy some punchy elements that make this you know people turn around and look so then you go diagnostic errors and you're like okay now now it's starting to sound like something treatment errors or now now it's really okay that's that's interesting because actually 16 diagnostic errors does sound quite punchy but it's very unclear what that actually means and you know like any any student listening any third year medical student listening will know that if you want to critique a paper go to the methods and you go well it it as you've written it rests entirely on the judgment of people reviewing emr notes and then there's another issue about ai marking its own homework as well so So yeah, I'll try and explain it, although there are other people with far deeper AI expertise than me who have run these types of studies before.

16:37So I think what they've built is a system that embeds itself in the EHR. And to be clear, the health system is very innovative. They engage with lots of quality improvement work. They've done lots of work with electronic tools in the past. So it's not like this is a health community that's totally new to the use of technology. These people are operators. They understand how this stuff works. The idea is that the EHR system is upgraded with a sort of cognitive companion, if you like, the co-pilot, to flag things as either red, yellow or green based upon whether or not the documentation has been properly done and seems to be going in the direction of good clinical management.

17:19The challenge is, how does one decide if a doctor seeing a patient has taken enough notes? if they have written down the right things if they have asked the right questions if they proposed the right tests um and the broader the diagnosis space the harder that is to pin down if i put 100 people in front of you with suspected appendicitis and showed them to 10 emergency room clinicians i'd expect a very high level of agreement about the questions and the urgency

17:48Dr Paul Wicks:and whatever yeah okay what if i made it what if we throw in fibromyalgia and diabetes and a brain tumor and the common cold and there's nothing wrong with you at all and you know and and and as that information space broadens and as we start adding other types of specialists okay psychiatrists emergency room doctors orthopedic surgeons and some from england and some from america etc and and this is what was done there was an independent panel of about 120 physicians from around the world only about a third of whom were from africa um asked to weigh in and to rate a subset of these cases as saying would they agree.

18:25And so the challenge is when you say, oh, there's been a 16 % improvement, that's based on the independent raters saying this is better or this is worse. But one of the challenges I identified in the results was that the raters didn't agree with each other brilliantly well. There wasn't that strong of an agreement. And so that's what I think puts the 16 % question in my mind. Because anytime you hear a declarative statement, like it's this much better you kind of go well actually that's an estimate isn't it it's an estimate and there's some range of error and what you hope is that some correct statistical method has been followed just to have some confidence that the true result lies within the range between you know we're saying 16 but maybe it's between is it between 10 and 20 or is it between minus two and plus but you know these are the types of questions you need to ask um the llm side was which was an interesting methodological application.

19:21But again, if we can't be sure of the foundation of what's going on, bringing in LLMs that, I mean, if I run the same query 10 times, I probably get 10 different answers, right? Like adding them as an additional marker. And in both the Microsoft paper and the OpenAI paper, there's a little hint of throwing shade at other models. So, you know, there was some stuff in the Microsoft AI paper where they tried to run the New England Journal of Medicine cases through other people's chatbots, and they sort of didn't perform as well. And you go, okay, well, so I have done those types of studies. And the challenge is you've got to be really careful.

19:57Like, what if the internet goes down? Do we reload it? What if it doesn't understand it? Do we help explain it? What if this one can take images and this one can't? So how much sort of spoon feeding you do for your model versus the competitors model is extremely important. And as you start drilling into a lot of this, it's kind of unclear. Well, who did the physicians work for? Did they work for Microsoft and OpenAI? Were they paid? Did they know what the overall thesis was? Because all of these things, again, there's no bad intent here, but all these things kind of introduce the biases that we know, like all of research, right?

20:33The replicability crisis of many sciences is through well-intentioned people adding up all these little cumulative heuristic biases. And so the whole process of guidelines and science and peer review has it's far from perfect but it exists so you try and iron those out it's really

20:52Dr Paul Wicks:interesting isn't it i've noticed in the in the sort of b2c world the consumer world the greater amount of discourse there is around certain health stuff the the more likely it is for people to start now quoting papers and people seem we seem to be in this era of well we don't want pseudoscience we want we want evidence but then i mean i don't even know the world of evidence to the level that you know evidence but i know the world of evidence more than joe blogs on the street because i'm a scientist and i'm trained in it but joe blogs on the street is also quite likely to go well i've read a paper and it said this and therefore that and actually we're in this world now where it seems for a lot of people evidence is binary they read this and they go 16 relative reduction in diagnostic errors chat gbt is amazing we should be doing this we should be doing and that's good will enough to remember you can remember that soundbite so when you're down the pub or on twitter you can remember that soundbite yeah and that that's pretty key here because actually the blog post one of the things that they wrote and you know i'm not sort of blaming the scientists here i worked in these types of companies i imagine the people that write the blog posts are not the scientists yeah but they compare the magnitude of this benefit um to the magnitude of some other things that have come through in medicine that also had a 16 percent same but they they talk about that as an effect size and that that's not really how effect sizes work So their quote was, these effect sizes of a 16 % improvement in diagnostic accuracy are comparable to antibiotic stewardship programs.

22:43Dr Paul Wicks:Whoa. Or alerts to encourage statin prescriptions in patients who need it. But this finding comes from a single system that can support a wide array of clinical decisions. And then they go forward and say, in absolute terms, the introduction of AR consult would avert diagnostic errors in 22 ,000 visits and treatment errors in 29 ,000 visits annually at Penda alone. So all the time here, so you've mentioned methods a couple of times. So there's the methods, there's the study that we've built, and there's the claims that rest upon this. And, you know, in the discussion, we sort of say more research is needed, right?

23:20So what we need is a bunch of evidence, usually meta-analyses and systematic reviews, ideally. These are the things that sit at the top of the pyramid of evidence, ideally done by an objective, neutral, independent group who does not have stock options in the company, to make these claims. But I do think it is puzzling why something like this could get through, because it's very clear it's not the case that this software simulation of green, yellow, red traffic light, is that going to have as big a magnitude in global health as statins and antibiotics? Really? Does this paper prove that point?

24:01Because if that paper doesn't prove this point, and this is just the first of many, what the community needs to be more vigilant of is are all future claims now going to need to get a lot more rigor, a lot more review before they've taken a face value? If we are indeed the intended customers of this information, are people with MDs and PhDs, are we meant to be critiquing this? Or are we such a small constituent part of the readership that maybe my deeply nerdy sleething is not important? Because we're seeing big announcements from the US AI strategy, right? We're seeing partnerships with the UK government about open AI and others having key roles.

24:42And I'm very far from a conspiracy theorist. but if i was a civil servant without scientific training if i was a politician without scientific training and i heard oh they're doing stuff in health they're doing stuff in kenya that sounds good maybe i i would be swayed by that um and so maybe we're not the intended audience here and that's why it's fine on some level that these rules and these norms have been violated because

25:06Dr Paul Wicks:this was never for us yeah and what comes to mind as well again playing like forcing myself here like to play devil's advocate a little bit is is to go you know is this a case of perfection is the enemy of progress and this is progress and actually we should get out the way of this because this is the future and it is going to be but like but then you know i'm i'm i'm clinically trained so then i go at like patient safety above everything else and i think if we start playing this out to well what's what what like you said at the start what is the potential damage here is that everyone is in the pub quoting each other going 16 relative reduction in diagnostic errors this should be doing our diagnosis and even i'm cringing saying the words because i feel like i i need a class three medical device certification on myself at this point to start saying this out loud but like yeah but if you look at things like the nhs 10-year plan it's saying these types of tools it is you're right yes so should we is this pre-print going to be used as evidence to show to somebody to go hey look we should do this and and then all the voices within the system do they have the time to to even check this right so so yeah so by the way i love gadgets there's a ai powered robot vacuum cleaner running around my kitchen right now so like love gadgets gadgets lifelong lifelong passion and but the perfect to the end of the good the thing that probably so good writing a second article when it risks being almost the same as the first article was that there is another study happening in parallel to the one that OpenAI talked about.

26:43So I'll back it up a second for anyone that's not deeply nerdy on evidence. There's all different types of studies one can do, right? So the lowest level of evidence is I'm Paul and I think this is great. That's expert opinion or face validity. It's basically meaningless, particularly if it's from me. Then you have case studies, right? I saw this one patient and they ate a carrot and their leg fell off. Carrot must cause leg fall off. Case series, a few more of those. You start going to objective studies. Oh, I've measured executives and they're all very tall. Therefore, there's something about being tall that makes you good about being executives.

27:13So just a cross-section in time, very prone to bias. Somewhere in the middle, you get to a randomized control trial where we split people out and they should be blinding so we don't know what groups people are in. And those could be small when it's a formative piece of research, or it could be very large when it's a big study with lots of variables we want to test. and then ideally we have lots of RCTs and then someone comes and reviews all the RCTs, adds them together, assesses them for bias and then we have this pyramid of evidence with systematic reviews at the top and low-level stuff at the bottom.

27:48Siri's just started spying on me so I'm worried now that Apple's...

27:53Dr Paul Wicks:Yeah that was Siri talking to Paul for those that didn't hear that in the background. so the open ai study is a retrospective observational study and it's called a quality improvement study which is a little odd for a couple of reasons one is that quality improvement studies often don't need ethics approval in many health systems i don't know if that's true in this exact case but usually what you do in a quality improvements survey is you go we're going to keep business as usual but we're going to make this change to the intake software or you know the flow of patients who get to see physiotherapy consults and we'll see what happens right so that and the aim is to get better interview you're not trying to prove a point in quality improvement studies you're generally not trying to say this is better than that because that's not the type of study it's equipped for at the same time as this study has been going on literally week by week there is a randomized control trial right by the gates foundation through an organization called path which has got a huge body of evidence and experience of working in lower middle-income countries with technology deployments and all the rest of it at the same clinics that study described its protocol in nature digital medicine late last year and published a full very detailed protocol with all the guidelines followed with all the proper literature review with all the statistical things and it was a lot more specific i think it's only looking at diabetes and hypertension so you know what i said earlier about that broad decision space and use the example of appendicitis well now imagine that this same software and it's the same software being tested in the rct was only looking at the accuracy of diabetes and hypertension now the good thing about diabetes and hypertension is we have tests for them that we have very high confidence about yeah whereas the problem with fibromyalgia or clown phobia is we don't have tests for them and so if someone agrees or disagrees in a diagnosis it's all a bit fuzzy Diabetes, pipe tension, we know if we're right or not right.

29:51So really well-designed, elegant RCT that's not mentioned at all in OpenAI's preprint, even though it shares at least one author in common. Not mentioned at all on the blog. At the very least, normally when you do this type of paper, you say, suture studies will explore blah, blah, blah. It was conspicuous by its absence. And the first thing that tipped me off was in the preprint, when you list the authors, You say, Bob works for Microsoft and Jeff works for the hospital and Susan works for a software company. It didn't have that. And that's important because I noticed that it didn't declare where the funding came from or where other sources of funding might have come from for the authors.

30:28And so when I saw this Gates-funded RCT, I expected to read, oh, we also have this nice grant from the Gates Foundation. Thanks very much. It seemed conspicuous by its absence.

30:41Dr Paul Wicks:I see. I see. So just making this absolutely clear then, the issue is that Bill Gates could just be puppeteering this whole thing. It's basically what you're saying. Let's clip that as a soundbite. The master stroke of the 5G chips in the vaccines. and so basically what i've concluded is this is all a massive play to get me to finally update to windows 11 windows 11 does not play roller tycoon 3 and so i'm out for that reason i'm out um uh yeah so i have no i have no personal thesis about why that is missing i'm just saying that i know if i did a formative piece of research and i've published lots of formative pieces of research to show you know it starts off as a c minus and then it's a c plus then it's a b and it's a b plus and then so when you say i've got an a grade product people believe you because they see you're working i see it is very odd that someone would publish a a type of study that's lower on the pyramid of evidence that makes broad claims um that don't seem to be supported by the evidence there that you know a preprint that doesn't conform to the usual structure of a medical paper when there's a really elegantly designed study that according to the original plan should be reporting out in q3 q4 so it's only speculation on my part here i have no inside information but it feels like something to do with timing i've driven why why july 2025 um and that's the part about which yes i like i said i have no interest in conspiracy theories um but um again if like with the microsoft thing it's a pattern if it becomes a pattern that when we need to tell a compelling story or something some bad pr has come out you know we've just laid off 10 000 people but here's a nice story about how in the future we'll be great at health that would be a problem or if we're trying to justify a very large investment that has been made from one company into another company and investors and others are saying, what did we invest all this money for?

33:02It would seem like a good time to go, look, great. We can help all of human health, which again, good aspirations, wonderful aspirations. That's what we all want. But if that becomes the driver of when things come out, we'd be in trouble. And I can tell you with absolute certainty, if you were a pharmaceutical company and you did that, you would be in heaps of trouble from large regulators with significant hours to compel you not to do that again. So there's a reason why when a drug company announces the results of its clinical trial at a meeting, that has gone through a very thorough process of vetting and review.

33:42Probably hundreds of people at the company, including many, many medics and scientists and regulators and all the rest of it, to ensure that the claim being made about weight loss and GLP-1s or what have you is robust. because otherwise if if biotech and pharma was a wild west we'd be in real trouble

34:02Dr Paul Wicks:100 and i've heard for so long now that algorithms need to be treated like drugs and it's so true like as soon as they're reframed in that way like it's so easy to just be like well why don't we do hold on why don't we do this why don't we do that like that's actually quite frightening like why aren't we treating algorithms like drugs like this is this is bizarre like why they can definitely have guardrails i mean go into chat gpt and ask it to make you a story about star wars using the ip and it'll say oh no sorry i can't do that yeah okay great um i've got this war on my finger is it cancer should i go to the doctor right now oh let me have a look now i'm a neuropsychologist if i looked at the water on your finger and tried to uh pretend i was a doctor i think that might be a crime so you know i i feel like it's um something we could at least look at and say all right well if it's going to go into doctor mode what else is happening differently in doctor mode uh sam altman's really recently said as a reminder to people who perhaps weren't already aware of it that your conversations with chat chibati are not privileged yeah you ask the legal advice from your lawyer that's privileged if we have a conversation about the aforementioned water my finger we have doctor-patient privilege.

35:13But if the thing is acting like a therapist, a doctor, a lawyer, and all that information could be subpoenaed, for example, then that leads us into some riskier territory for patients. So at a minimum, I mean, getting a little bit off track, but I once looked after a database of 8 ,000 people with HIV, some of whom were in countries where it was illegal to be gay. And so the potential could be that this type of data, if leaked or subpoenaed or what have you could end up you know in in terrible consequences for for people in in authoritarian regimes and others and you know um there's a reason medicine frontier does a lot of their work on paper indeed because when you create scalable systems also creating some significant vulnerability so anyway that would be down the more yeah do me do me end of the scale yeah i think these i think the thing that's crazy is these things are all addressable microsoft and

36:10Dr Paul Wicks:OpenAI have the smartest people on the planet fighting each other tooth and nail to go and work for them. They have$50 million programs of research with Stanford and Harvard and all the big places. They have access to the brightest minds of a generation. So this is all eminently addressable. I just think we need to kind of maybe draw their attention to a bit of focus and kind of go, okay, you can't do move fast and break things in every domain. Indeed. um and you know i do believe there's a community out there willing to work with them and collaborate and support and and you know uh lay out why the guidelines are there and be quick and you know move fast and and if so then we'll make the trust but you know you only need to go far as back as ibm watson health to look at cases where steam rollering with big press releases that don't hold water will only last you a year or two and then and then you're a bust and and you know reputation has popped and you can't get it back again linked to harm you mentioned harm and geography there in the case of hiv um and a lot of programs medicines on frontiers etc um operating in low middle income countries do you do you read anything into this being done in the global south so i really only have the evidence from the manuscript to go from uh and the rct uh that's happening in parallel and what i see from that is that this group approached open ai and a doctor there who's very into technology and said we would like to try this so yeah i mean that's the that's the only information we have um certainly i think we do have obligations um in all sorts of ways you know we're based in the uk you know we're a country with a long history of being colonizers and we should be aware and conscious of the influence we've had on other health systems other education systems and and many of those issues continue to affect communities well well into the the present and beyond um but yes i suppose one would have questioned about the level of ethical oversight that is available in different countries or perhaps the enthusiasm that might go unchecked in a lower resource setting of the potential for unequal potential benefits of working with a large organization like this.

Read the full transcript

38:39If you're a very prestigious institution in America or the UK, it's pretty commonplace for you to work with big tech. You have very good partnerships. Perhaps if you're an LMIC, maybe that balance of power is a little bit different. And so you might think we should be extra super-duper attentive to that. Because I think, particularly with the risk that so much of AI is trained on white, middle-class, Anglo-American folk who look like me, that might be a place where, in particular, issues like algorithmic drift could be of greater risk to populations who don't look like me. um and so you know at the moment we're talking about relatively simple things about you know did you write this note properly but if in the not too distant future it starts to look at things like um genetics you know um so in the nhs 10-year plan supposed to be whole genome sequencing lots of people i know for a fact there are not enough genetic counselors if someone's trying to return genetic information back to people you know that that's kind of a risk there um you know we uh we also know that different groups are affected differently by different conditions different medications and you know we just probably need to show our working to kind of go you know how have we taken this into account so one one sort of an annex data that i put in there is um about kenya in particular i happen to know this because i did a piece of uh of writing about some research that took place in south africa um looking at the epidemiology of different and countries in uh in africa and um was interested that there'd never been a confirmed case of ebola in kenya so when the preprint was saying we took into account the epidemiology and the clinical guidelines of kenya i just sort of posed this example so so what does that mean does that mean that the algorithm would know that there's never been a case of Ebola and so discount that which might be okay because otherwise everybody shows up with a nosebleed you're going to potentially call the quarantine team but what happens if someone does what happens if you get the first case what if someone comes in from neighboring country and is infected if the algorithms play to downweight it because the epidemiology great how do you how do you then adjust to situations on the ground say if a mysterious new illness appeared from wuhan china you know so you know we've we've just been through this everybody it seems like you know we have we do have the time that we do have the resources we do have the experience to incorporate some of this knowledge into the systems that we use but we have to progress systematically incremental advances you know radical incrementalism i think will win the day here and build systems that we all trust with radical transparency um absolutely

41:18Dr Paul Wicks:i mean plus or minus radicals but perfectly honest just transparency i think that that seems to be the theme running through this for me and it and for whatever reason this keeps coming up with ai and llms particularly trust and transparency it keeps being an issue um product wise it's obviously the lack of determinism the black bottle like all of that stuff can lead to this lack of trust and and it's it's not a transparent system i can't see the innards and therefore that's perhaps where we start but if then you layer in the lack of transparency associated with essentially press releasing a pre-print as if it were perfect evidence if people like yourself are going to call this out we are going to appropriately start to lose trust in it because the question for anyone sensible or you know scientifically minded will then go well why couldn't you wait why couldn't you why couldn't have microsoft waited last time why couldn't have open ai waited this time and that's a genuine question i don't i don't sit here wanting to cast judgment on either organization for fear of disappearing um i would love these things to work i would love these things this is it i would i would love to put in weird symptoms at 2 a.m into a system that is not making assumptions or is judging or is sleep deprived or this you know on his third shift or whatever yeah i want these systems to work but i'm going to put my health and the health of my family into them that's it so i have questions and trust me bro isn't cut in it I've seen it not work enough times that I shall not trust but verify is the process there's a reason all the B2 stealth bombers are not in hangars it's so that the Russian satellites can see them so that they know how many nukes the Americans have that are airborne and vice versa that's why the Ukrainian drones could blow up all those TU-95s because trust but verify means show you're working if we're doing that for the existential nuclear threat of mutual mutual destruction if ai is that great and that powerful and it will become that ubiquitous i think we need to to treat it in the same way um show me your argument for founders listening who might be i don't know about to pilot an ai tool in healthcare what would be your single best piece of advice on evidence generation you mentioned from trust me bro right up to meta analysis this is the this is you know i mean this question actually has been going on for ever since i've been in health tech i've people have been trying to figure this out how much evidence is appropriate but an ai tool in healthcare specifically what would you what advice are you giving to them so in the uk yes in the uk there's a group called circe ai with a ton of expertise from the clinical, the scientific, the regulatory side that exists to facilitate this type of work.

44:29And then there will be other local groups in the US and other places. So I think it's a real high priority to work with those groups. For the evidence, so this is a personal little windmill of mine. I think there's kind of a bit of an N-shaped curve, right? Like a dose benefit curve. You need to have some evidence. you cannot spend your entire you know five million dollar seed round doing nothing but producing papers that would be bad that would mean you're a computer science lab right so there's some sort of relationship between how big a claim you're making what's the scale of what you're trying to do and what can you do the good news is i think you can start out small publish your user experience research publish your formative hey we tried this and it didn't work start with posters go to conferences, engage with the community, and you can kind of build it up from there, really.

45:20So, yeah, I don't think it has to be, you know, one perfect RCT in the Lancet proves everything. You build up a portfolio. So the Digital Medicine Society has got some really excellent work in this area. I worked on something called the Evidence-Defined Criteria with a payer, a US payer, Jordan Sutherland and he kind of mapped out here are the things that increase our interest our credibility of the evidence here are the things that decrease it and it was specifically designed for digital health to take away some of the flags of oh well professor x from university y endorsed it that actually melts you down on the evidence defined criteria so yeah and then you know um I think you can you can work with a scientific advisory board and really sit with them and go well what evidence

46:04Dr Paul Wicks:would persuade you to deploy this? If the main proponent was your academic enemy, what would be such persuasive evidence that you would go, okay, well, I'll give this a shot, right? Because in at least hospital medicine, things like WhatsApp, things like scribes pour through the system with no need for RFPs because they just worked. You could see them just work. That's not going to be the case for this. This is going to be large complex.

46:37and iterative. And it is a little bit different for areas like radiology, dermatology, mental health. We see a lot of these things. So they each have slightly different aspects to it. And I think particularly with their medical devices, it has to be in lockstep with a regulatory pace. So yeah, I think sometimes maybe that's what puts you in lockstep of, I have to prove something to a regulator. I might as well prove it to the market. And it can also be a good spurt to build your IP strategy that way as well. Once you've declared something as the state of the art, it can affect your timeline for filing pants and things like this.

47:10So yeah, it's an investment. We talk about your proof stack. You have a tech stack, you need a proof stack. And you can't run, most startups probably working on about 40 pieces of software to run their business. I think it's the same with evidence, right? There's some expert opinion, there's some advisory boards, there's some trials, there's some studies, there's some surveys, there's some health economics. and it builds and really your sort of credibility is is the sum of your proof stack much like your technical excellence is the sum of your tech stack i love that thank you paul um it's a really nice

47:44Dr Paul Wicks:framework that i definitely yeah i'm definitely going to commit that to memory and start repeating that and i will quote you um i'm left uneasy about this if i'm honest um it's i'd love to actually get someone from microsoft on to talk about this i'd love to get someone from open ai on actually to talk about this so if anyone is listening that knows anybody in either of these teams that wants to ping this to them and ask them to come on for me or do an intro then please feel free but i am left uneasy and i think it's because like you when you're writing this you sort of know don't you sort of you'd expect whoever's written this and someone in the chain of the teams that have gone from writing it to to press releasing it and publishing it someone or a few people would know that hey we at least should put by the way this is just a pre-print right at the top and make this clear and that hasn't been done and i'm really intrigued as to why i'm really intrigued as to why why that level of just simple i i don't know what you'd call it academic decorum or or like following clear principle or following guidelines set by people that make the rules like why what why hasn't that been done that's the question that i have here and i think that that's rooted in a desire for above all else patient safety and neither you nor i paul want to be in the way of progress that's not why we do what we do but you and i as clinicians will will just be so trained and conditioned to to know that that is only ever appropriate when patient safety is put first and to lord claims in front of people that may or may not be statistically significant should feel abhorrent like it should it should for any for anyone that has that level of integrity required to publish this should should not feel like that is okay and that's what i'm really intrigued to get to the bottom of here so again just going to say it would love someone from either of those organizations to come on and talk about this um but that's kind of where i sit what would you want to leave people with paul that are listening to this conversation be it from those organizations or be it from the health tech industry more broadly um my one by the way is a pre-print is not a paper that's the that's what i would love everyone to take from this yeah so i think um look i think i i predict that there were voices within each of the company that said this.

50:40And sometimes it speaks to whether or not scientific and medical voices are to be consulted or if they actually have authority within a company. And I think if you have a chief medical officer, if you have a chief scientific officer, if they pull the red cord and it doesn't do anything, then you have a problem there. So I think that that's the case. I do think that we need to ask ourselves when a big tech company who, through pension funds and and shareholdings like, on some level we own, and we buy their products and we're their consumers. Who controls who here? Do we need to be grateful that these organisations have deigned to drop a crumb of their evidence?

51:21Or actually, are we the people who control these systems? Is there a risk that these things are, in some ways, anti-democratic, and that actually there's a more strident view that we need to take? And again, I think there's plenty of opportunity to reset that dialogue and understand, you know, where is a benign influence and well-practiced and what have you, and where are people chasing hype cycles. And I think the degree to which we understand that and that we believe in that as a community will only accelerate the growth of these companies and speeding up of all the things we want to do, faster diagnosis, better diagnosis, faster treatment, a more personalised care.

51:59But yeah, the thing I'd leave with is my favourite Carl Sagan quote, So extraordinary claims require extraordinary evidence.

52:08Dr Paul Wicks:I'm just going to end it there. That was glorious. Thank you, Paul. Perfect. Great fun. I have to run to a client call.

From the publisher

When OpenAI and Penda Health dropped a preprint claiming their “AI Consult” tool cut diagnostic errors by 16% in Kenyan clinics, headlines lit up across healthtech. But veteran digital‑health researcher Dr Paul Wicks spotted the cracks.

00:00 Introduction to OpenAI's Preprint and Its Significance

05:55 Understanding Preprints and Their Impact on Research

11:58 Evaluating the Methodology and Results of the Study

18:01 The Importance of Trust and Transparency in AI Research

23:49 Conclusion and Final Thoughts on AI's Role in Medicine

29:50 Randomized Control Trials and Their Importance

36:47 Building Trust in AI Systems

43:05 The Need for Transparency in AI Development

More from The Healthtech Podcast

All 65 episodes
#408 First Look: Open AI cuts diagnostic errors by 16%. Or does it?The Healthtech Podcast · 52 min
Listen in VO