In short
Eye On A.I. - Episode #185 Summary
Episode Title
Damon Rasheed & Saurabh Jain: Achieving Up to 90% Accuracy in Predicting Drug Clinical Trial Outcomes with Opyl's TrialKey.ai
Host
Craig S. Smith
Episode Overview
This episode features Saurabh Jain and Damon Rasheed from Opyl discussing their innovative AI model, TrialKey.ai, which can predict clinical trial outcomes with an impressive accuracy of up to 90%. The conversation delves into the methodologies used, the challenges faced in clinical trial design, and the implications of these advancements for the pharmaceutical industry.
Key Topics Discussed
- Introduction to TrialKey.ai
- Prediction of clinical trial outcomes with 90% accuracy.
- Utilization of a dataset comprising over 400,000 past clinical trials.
- Focus on extracting relevant variables for trial predictions and optimization.
- Challenges in Clinical Trials
- Poor trial design: Approximately 60% of trials are poorly structured, leading to failures.
- Need for better methodologies to enhance trial outcomes and save resources.
- Data Processing Techniques
- Use of Natural Language Processing (NLP) to extract critical variables from complex datasets.
- Identification of hidden factors that can improve trial success rates.
- Implications of AI in Clinical Trials
- Revolutionizing trial design for better predictive capabilities.
- Optimization of resource allocation for pharmaceutical companies.
- Enhancing decision-making for investors by providing insights into trial success probabilities.
- Practical Applications and User Interface
- The user-friendly interface allows companies to input trial data and receive predictions.
- Features a clinical trial simulator to model potential outcomes based on various inputs.
- Future developments may include optimizing trial variables through a slider feature.
- Market Potential
- Potential for a vast market as the number of annual clinical trials could increase dramatically through effective utilization of AI.
- Prediction of a shift towards AI-driven trials, reducing the need for physical human trials.
- Future Developments
- Plans to incorporate data on FDA approval processes and aftermarket performance.
- Exploration of other applications of their modeling approach in various sectors, including sports and digital marketing.
Expert Insights
- Damon Rasheed (CTO) shared his background in sports betting analytics and highlighted the parallels between predicting clinical trial outcomes and sports outcomes.
- Saurabh Jain (CEO) discussed the foundational aspects of TrialKey and the significance of utilizing big data in clinical research.
Conclusion
The episode emphasizes the transformative potential of AI in the pharmaceutical industry, particularly in enhancing clinical trial success rates and optimizing resource allocation. The conversation ends with an encouragement for listeners to stay informed about the evolving landscape of AI in healthcare.
Call to Action
Listeners are urged to follow the latest discussions and innovations in AI and healthcare, subscribe to the podcast, and explore the TrialKey platform for insights into clinical trials.
---
Additional Links
- [Twitter: Craig Smith](https://twitter.com/craigss)
- [Twitter: Eye on A.I.](https://twitter.com/EyeOn_AI)
- [TrialKey Website](https://opyl.ai)
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00So right now the model that we have can predict the outcome of a clinical trial with about 90 % accuracy. Everything about how humanity interacts with drugs and devices and alternative therapies is actually known within that data set of 400 ,000 trials that have existed in the world. And all we're doing is pulling out the relevant variables for every phase, for every condition, and just putting it together, and then using that as a knowledge set to make future predictions. From what I've seen so far from TrialKey, it's about 60 % of the world's trials are designed poorly. There's major design flaws in them.
0:32So I think that there's a lot of drugs or devices that would probably have been successful had the trial had been structured correctly. It's not the underlying molecule or compound that was the problem. It was the actual design of the trial that led it to failure. So I think hidden in the data, there would be thousands of these drugs or devices that would have shown efficacy. Hi, I'm Craig Smith, and this is Eye on AI. In this episode, I talk to the people behind Opal AI, that's O-P-Y-L dot A-I, an AI system that predicts with remarkable accuracy whether or not a clinical trial will be successful.
1:16Saurabh Mishra, CEO of Opal, and Damon Rashid, CTO, explain how they use natural language processing to extract key variables from trial data and machine learning to predict outcomes. By analyzing over 400 ,000 past trials, their model can forecast success rates for new trials, helping pharmaceutical companies optimize trial design and helping investors make informed decisions. I hope you find the conversation as amazing as I did. Hi, I wanted to jump in and give a shout out to our sponsor, NetSuite by Oracle. A little quick math. The less your business spends on operations, on multiple systems, on delivering your product or service, the more margin you have and the more money you keep.
2:10But with higher expenses on materials, employees, distribution, and borrowing, everything costs more. So to reduce costs and headaches, smart businesses are graduating to NetSuite by Oracle. NetSuite is the number one cloud financial system, bringing accounting, financial management, inventory, HR into one platform and one source of truth. With NetSuite, you reduce IT costs because NetSuite lives in the cloud with no hardware required, accessed from anywhere. You cut the cost of maintaining multiple systems. You improve efficiency by bringing all of your business processes into one platform, slashing manual tasks and errors.
3:00Over 37 ,000 companies have already made the move. So do the math. See how you'll profit with NetSuite. By popular demand, NetSuite has extended its one-of-a-kind, flexible financing program for a few more weeks. head to netsuite.com slash IonAI. That's IonAI, E-Y-E-O-N-A-I, all run together. Again, head to netsuite.com slash IonAI, netsuite.com slash IonAI, for its one-of-a-kind flexible financing program. So Damon, why don't you go ahead, introduce yourself, and then Saurabh, you can you can go next hi yes i'm damon rashid uh my background is i'm actually an economist so i spent the first few uh years of my life working for the equivalent of the department of justice over in the us in australia it's called the accc so federal regulator um but in the background i was also i've got a always had a fascination with data so i've been running sports betting models models, horse racing models, greyhound models, trying to predict the outcome in these competitive sectors, trying to get an edge over the bookies.
4:20So my day job was working as an economist at the regulator. And by night I was trying to run models to beat sports betting and racing markets. And these days I run a data science team and we specialize in providing machine learning solutions, I guess, where there's difficult problems to solve. We do a lot of consulting for banks, insurance companies, universities, road authorities. So it's been a fascinating journey. Yeah. And Saurabh, why don't you go ahead? Yeah, totally. So I'm a software engineer by trade. Back when I did my thesis like 20, 25 years ago, it was actually machine learning. It was all about how do you predict the outcome of patients getting cancer based on about 50 odd variables.
5:08And back on those days, you had to literally hand code everything. So I did my 70 ,000 lines of code to get that model working. Then after I did a tech startup, exited to Telstra, which is like the equivalent, like an AT &T here in Australia, got into sales for about a decade, got into management, was the CEO of a public listed tech business, grew that from about 20 mil to about 100 mil, exited that about three years ago and then stepped in as the CEO of TrialKey about six months ago now. TrialKey is the new name for Opal, is that right? Or is Opal the new name for TrialKey? So Opal's the company name and that's what we're on the stock exchange at.
5:49So ticker code is OPL. The actual product that we take to market is TrialKey. I see. And Damon, you're the CTO for TrialKey, is that right? Yeah, that's right. and we've been working on trial key for about three years now um and we've only just got it to market in the last probably the last four weeks so this is the exciting period where we can uh start to get some feedback about our product yeah uh do you mind starting damon by talking about the sports betting uh which i find fascinating i i told you earlier i i uh had a turn at horse racing prediction prediction using uh you know supervised learning classifier what kind of models were you using uh how what was the success rate uh and then how can you relate that to what you're doing at trial key yeah good question so this is around 2005 that uh you know started building sports models and mainly for u.s sports um and some international sports like tennis and and golf but we were doing um national hockey league uh nba um you know nfl um so you know the prediction models at that stage there wasn't a lot of competition there weren't a lot of i guess statisticians modeling sports back in 2005 so there was quite an edge over the bookmakers so you could take publicly available data you could analyze it using pretty simple models um back in day like you know traditional linear proba or logit models um you know even some simple logistic regression uh provided some uh some advantage over the market but as time went on um you know more and more statisticians came into the market data became more readily available and the market became tighter and tighter and it became harder and harder to win so you needed more advanced models and um you know at that stage you know I had other things on so sports betting sort of took a back seat and um and then we started working on this this trial key product um so yeah and I what the way that I see it is that you know predicting the outcome of clinical trials which we'll talk about in a second it's kind of like sports betting 20 years ago there aren't a lot of people doing it at the moment.
8:19It's a very virgin industry and there are opportunities here. It's brand new. And that surprises me given the amount of money spent on clinical trials. That's right. I haven't heard anybody doing this. I mean, Sir Rob and I had a conversation earlier about Insilico, a company that I've had on the podcast and that I follow somewhat. And And they're focused on narrowing, you know, increasing the chances of a successful clinical trial by narrowing the scope of small molecules that they're going to put into trial. But I don't think they're doing any predictions on the clinical trials themselves. Is there anybody doing this that you're aware of?
9:12So there's a couple. I mean, I think Insilio are kind of playing in this area, but you're right. They're a lot more on the molecule side. There's another company that we've seen out of Israel that are talking about doing something, but they haven't launched something to market yet. So I think right now we have first mover advantage, but it'll be like the sports betting in a couple of years. Everybody will be doing this. Yeah. And what happens, Damon, with the sports betting? The odds, the more people that do it, it narrows the odds. So the payoff of a good prediction is fewer and farther between.
9:51Or is there something else that flattens the curve and makes the opportunity less attractive? No, I think you nailed it there. So it's more and more people, you know, having predictions and then moving the market, but also the bookmakers themselves have invested significantly in data science capabilities. I know one bookmaker, one of the biggest that has 50 data scientists working for them, making predictions. And also a lot of professional outfits have what I call private data, where where you have this publicly available data, but for some sports like tennis and soccer, those publicly available data sets aren't that detailed.
10:40So they have people watching these events and coding their own variables and at scale. And I know, particularly in the UK, some very successful outfits that have won literally billions of dollars by creating their own data set, ignoring what's publicly available and actually hiring resources to create their own data set yeah uh and and sir rob you were talking when we spoke earlier about investing in clinical trials as as an investor i mean investing in in these small pharma companies that uh that are are you know at very low uh i don't know what the term is but their their their cost of capital is uh pretty high in there uh yeah and so you can get in uh fairly inexpensively and if you predict a successful clinical clinical trial and they are successful then the value of the company uh surges but and i can see maybe that opportunity would narrow as it did in sports betting as this kind of thing becomes more prevalent.
11:58But the value of predicting a successful clinical trial to the companies remains the same, right? I mean, it saves them time and money if they know which drug and which protocol is likely to be successful. Is that right? Yeah. So, I mean, just even a bit more context. So right now the model that we have can predict the outcome of a clinical trial with about 90 % accuracy. So we get it right most of the time. So if we think about the use cases, I think totally one is on the investment side. Because often the probability of success is actually not baked into the share price for a lot of these public companies.
12:46Damon and I were one for one. We invested in a company called Dumerix. We got in at a 10 or 12 cents about two months ago. They announced a successful trial mid-March. Now it's about 30 cents. We've done quite well out of that. We probably need to do this 10 more times just to make sure we didn't get lucky once. But I think the fundamental theories for one or two trial, one or two drug company, either a success or a failure will have a dramatic impact on their share price. So that's one use case. And the other use case is for the drug company themselves. So if they're running a trial, they might have two or three different options that they're doing.
13:19So we can help them figure out what's most likely going and succeed. And because it's explainable AI, we'll tell them what variables to tweak. So for example, they might have 700 patients, but they actually need 900 for this phase, for this condition. Or they might be using dosage of 25 milligrams, but the data shows for this condition, 50 milligrams is better. or they might be doing a cream, but the data says, look, a nasal spray is better. So all these kind of nuancy things that we try to do, we did this similarly for another company recently, and we kind of increased their chance of success by about 30%.
13:52And this is all stuff that's hidden in the data, but just people don't know because the data set has just gotten so big over time. Yeah. And let's back up a little bit. I'm sorry I jumped right in on sports betting. But so drug companies, they spend a lot of money in hunting for a pharmacologically active molecule or compound. And then they go through various animal trials. And then they're, at least in the United States, I think it's three phases of clinical trials. and that whole pipeline is very slow and very expensive. So you don't want to spend money on a clinical trial and have it fail. You want to pick your bets very carefully.
14:47And that's what trial key does, right? I mean, that's the value proposition. And can you tell us what kind of data are you using? what kind of data do you need from companies to reach that 90 % accuracy that you're talking about? Totally. So a bit of context about how we got the data source. So what Damon and the guys did, they've sucked down about 400 ,000 clinical trials from clinicaltrials.gov. So all this data is publicly available. But what took the three years of effort was to be able to scrape 700 variables off each of those trials. And three months ago is about 500. A year ago was probably about two or 300 and we're adding more and more data sets from those variables and I'm sure Damon can talk about some of the LLMs that he's kind of worked on to help do that but then the fundamental thesis we have is everything about how humanity interacts with drugs and devices and alternative therapies is actually known within that data set of 400 ,000 trials that have existed in the world and all we're doing is pulling out the relevant variables for every phase for every condition and just putting it together and then using that as a knowledge set to, you know, to make future predictions.
16:00I know that's a good summary. So, so that the challenge that we've found was that the actual publicly available data set clinicaltrials.gov where most, you know, significant trials are registered is a, is a really badly curated data set. I was going to say more nasty things about it but that's probably the um and the problem is that there's no real onus for pharmaceutical companies um drug development companies etc to report into that database so it turns out that only about 17 of all trials ever done we actually know what the outcome is that they've been officially reported um which is when you think about it that's that's a failure um you know, it's a data collection failure because as Zyrab said, if we know those, you know, the results of those trials, it can help in the development of new drugs, new medicines, new life-saving interventions.
16:57But because the data is so badly curated, we've spent three years trying to fix that problem and actually generate our own data set based off that initial publicly available data so that's what's taken the time and part of it is around identifying which trials actually passed and which ones failed but also in the free text that's written generating a whole stack of variables using advancements in natural language processing as Saurab mentioned that technology just wasn't available two or three years ago with you know chat GPT etc we can extract a huge amount of data from those clinical trials, the free text within those clinical trials using language models.
17:43And I don't think anyone in the world has ever done this before. So we're pioneering it and the results and the insights have been fascinating. So when we go out and talk to people in the industry, they're totally engaged with the results that we're providing because they've never seen this sort of information. They've speculated on it, but they've never seen it. When you say that part of the problem is identifying which trials are successful, aren't trials publicly announced, I mean, at least in the United States by the FDA? So once you've done your three phases, then you can apply for FDA approval.
18:23If you never get to that stage, let's say you have a phase two trial and it doesn't succeed, there's an obligation to report but but it's not really police so you don't know the status of that phase two trial and the trials aren't linked either a phase three trial you don't know what the corresponding phase two trial was there's no in the data set that just doesn't exist and initially we thought if some if a trial wasn't reported on we just assumed that it must have failed but that's not actually the case at all there's many many trials thousands of trials that actually succeeded but the data was just never updated in in the public source so you know that that's um that made our job difficult but um but it was rewarding to sort of solve that problem as well yeah and the a certain amount of this data that that you can get publicly i would guess is in tabular form already uh is that right what percentage of the data comes to you uh structured and then uh i want to ask about how how you extract data from you know unstructured text yeah yeah so out of the 700 variables we've got probably about 300 are structured and that's things like you know how many patients they're looking to recruit or what countries are they in and um you know who's involved in the trial that sort of information comes structured there's a lot of structured data but then the unstructured the data is the fascinating data and we're talking about things like extracting you know the mechanism of action for a drug or the actual molecules and compounds themselves or it might be you know dosage levels or information about the inclusion exclusion criteria on a clinical trial which is you know who could get selected or or who can who misses out on the trial even um but the end points that they're trying to reach their primary and secondary goals of the trial.
20:17All of this information is free text and the only way that you can gain insights from that data is to create that free, take that free text, apply a natural language processing model and turn it into more structured data that then you can run through a model. So what we've been doing is creating the extra 400 or so variables which will increase significantly in the next six months. We've been using language models to create those variables from the free text. And the insights from that, it's been fascinating. So we can, something as simple as once a patient's enrolled in a trial, how many touch points they have with a particular site or hospital, that's critical information as to whether a trial will succeed or fail.
21:05And those insights were unknown before products like TrialKey came came onto the market. And that's very, uh, nuancey, like a variable like that. Um, if you're doing like life-saving cancer treatment, you tend to not need to touch the patients a lot. Um, but if you're doing a lot of alternative therapies, a vitamin-based observational study, you tend to need to engage with the patients like every couple of weeks. Um, and that's the real kind of nuancey stuff that the, that the models kind of picked up on. Yeah. And, uh, to extract the data, um, I mean, I imagine a lot of this is, at this point, prompt engineering.
21:43Is that right? Where you're working with a large language model on a block of text, and you're asking it to, first of all, surface all the variables that it can find in the text. Is that right? And then you decide which of those are likely to be important. Yeah, so yeah, exactly right. So the prompt engineering is a massive part of what we do. So, you know, we're speaking with subject matter experts and they tell us what they think is important around a trial. That enabled us to create some prompting to try to get that information from the unstructured data. and then to do that on mass like to do it for 400 000 trials um and you know lots of prompting as well to extract as much information as we can it's a big exercise it takes about a month for our servers to run through and collect that data on 400 000 trials and then rather than um us assign weights to each of those variables that are created we create a model and then the machine learning model determines whether that variable has any impact on the probability of success or not.
22:59So the machine learning model applies the weights. So for a lot of the data we collected, like a very, it had a very minimal impact, but then, but some of the data is, you know, those variables are incredibly predictive. Wow. And it's not to the point yet where you You can simply ask a large language model to look at a drug and a protocol and predict based on the weights in its memory whether or not that's likely to be successful. I mean, you're using the large language model to extract parameters, and then you're putting those parameters into a structured format and using a classifier, essentially, to decide the probability of success.
23:53Is that right? That's exactly right. We actually did try what you just suggested as well, and just to see if the natural language models could predict if that was the prompt. but the predictions that they came out with you know very well correlated to our our model so i don't think they're at that stage yet where um they can actually add value to the predictions over and above what our model uh what what our model does i don't i just don't think it's um it's in it's in its realm of capabilities just yet maybe one day yeah uh and then do you also So in the training phase, do you feed post-launch data into the models?
24:45Because even if a clinical trial is successful, the drug may not be successful in the marketplace. Do you use any of that data? Not yet. That's something that's on our development horizon. And also working out the probability of FDA approval as well. Those two things are in our development horizon. We've spent most of our time so far predicting phases one through three of clinical trials and haven't looked at a lot of aftermarket data. but yeah that's definitely a use case for the future. For now we're predicting something very narrow we're predicting whether a trial meets its primary endpoint which is really does it answer the question that it's set out to solve so we're not really predicting whether it's going to make money or whether it's going to be successful.
25:39The other thing that we do use the predictions for is a competitor analysis so remember that Domerix trade that we told you about that would serve a kidney disease. So their probability of success was about 55%, but they had about a dozen competitors trying to solve the same problem. But most of those guys had their probabilities in like the teens, 10, 15, 20 % success. So that's why we invested in them. We saw them as like a real outlier. Then the other part that model predicts is when they will succeed. So you get a nice kind of, almost like a Gantt chart that says, here are you or your competitors and these ones we think will succeed and these ones we think will fail.
26:16So it gives you a bit of a market analysis, but it is still very, very narrow around the primary endpoint. Yeah. And the primary input, can you explain that a little bit? Yeah. So for every trial, there's a goal that they're trying to reach and they're called, you know, either primary or secondary endpoints or hypotheses. And sometimes or often there's more than one primary endpoint and it might be, you know, is this drug that we're testing in layman's terms, does it outperform what's currently the benchmark, you know, for this particular condition? Or it might be an equivalence test. It might be, does it perform at least as well?
26:59There's all sorts of, you know, different primary endpoints and there might be secondary endpoints as well. So, but our model, we set a very clear target for it to predict, and that's whether a trial will meet one or more of its primary endpoints in the phase that it's in. But there are other targets that our model could look at in the future, and we've done some experimentation around this. And it could be, is the trial likely to complete on time? Or is the trial likely to reach its target level of patients? You know, these are important questions as well that, you know, drug designers would like to know from the outset.
27:39You know, like how realistic is it that we could get 700 patients for this condition with this number of sites and this number of countries? Or how realistic is it that we can complete it in 12 months? So there are other goals or targets that we could predict. But at the moment, we're focused on whether a trial would meet its primary endpoint in the phase that it's in. So, primary endpoint doesn't necessarily, reaching the primary endpoint doesn't necessarily correlate with a successful trial. Is that right? Usually if it meets its primary endpoint, it can progress to the next phase. But there's commercial decisions at play as well, right?
28:22It might be whether it smashed it out of the park or whether it just met its endpoint or whether competitors have come in while this trial has been running. There's lots of other factors as well, but it doesn't necessarily mean that, you know, a drug's going to go all the way through to FDA approval. It just means that it's eligible to get to the next phase if it's met one of its primary endpoints. Right. And does next phase mean phase one clinical trial, phase two clinical trial? Is that the phase you're talking about? Yeah, so phase one is usually around, you know, is it is something safe to use?
29:01And then phase two, it's more around testing safety, but then you start to test efficacy as well, whether it works. And then phase three is testing whether it works on a much larger sample of patients. So once you progress through all those phases, and you know, successfully, then you that's when you can apply for say FDA approval, you know, of simple that a bit, but that's generally the journey. Yeah. And right now you, you're offering it for meeting the, the, the phase one end point. Is that the, the, where the product stands right now? No, we cover all phases. So we could help design any trial that's in phase one, two, or three.
29:49So yeah, we cover all those phases. The predictions are pretty accurate across those three phases. Phase two tends to be the hardest for, you know, to get a successful trial, because that's where the rubber hits the road and you're trying to get efficacy for the first time, right? So if you're running a phase two trial, the probability of success is usually a bit lower than if you're trying to test safety or in a phase three, trying to prove efficacy on a larger sample. In your test set, I mean, as you're developing this, this 90 % accuracy, Sareb, that you're talking about, is that across all three phases of clinical trials?
30:32Have you tried predicting FDA approval or do you need more data or you've tried and your accuracy isn't where you want it to be yet before you start offering that? Yeah, yeah. So the accuracy between phases is pretty consistent. And also, interestingly, for rare or novel conditions, the model predicts well in those areas as well. There was, you know, some concern that if something's never been tested before, a drug's never been tested for a particular condition, how can a model predict success rates? but it turns out that there's enough data on, you know, how trials are run in generally, what are good practices that are beneficial for rare and novel conditions as well.
31:26So I actually do see that products like TrialKey might actually lead to more rare conditions being investigated and researched because you can have greater confidence of outcomes, as a side note. So in terms of FDA approval, that's probably one of the next models that we're going to tackle. We're trying to get as much data as we can around FDA approvals. It's a bit, yeah, the data sources aren't as transparent as say clinicaltrials.gov, where everything's registered. So we know what comes out the other end if something is approved by the FDA, but it's a little bit opaque about what actually got submitted to the FDA.
32:10So you can make some guesses about if something passed the phase three. Yes, it probably did go through that process, but may have failed. So there's a few data challenges there, but something that we think we could get a model for in time. How different does the model behave when looking at U.S. clinical trials or European clinical trials? or I don't know which protocol Australia follows, whether it's closer to Europe or the US. Is there a difference or do you have to tweak the model for the regulatory environment that it's being tested in? Great question. So it turns out that there's a number of countries around the world that have their own database where clinical trials can get registered, but most significant ones will be registered on the global one, which is the American one.
33:12In terms of the way trials are run, there's enough consistency in the parameters between different countries that we can match up the data set. So for example, in Australia, we have a database called the ANZCT, but it's almost identical to the one in the US. So we can map that data so that we can combine those data sets. And there's other ones around the world as well. And they're run reasonably consistently. What is interesting, though, is that some countries systematically overperform or underperform for various conditions or phases. And that's one of the insights from TrialKey. If you're running a particular trial for a drug and a condition, that there's countries and sites that are going to be beneficial to be in and ones that aren't going to be beneficial and it varies between phase and condition so it's incredible insights there where you can significantly improve your success rate just by knowing what jurisdictions to go to and it might be something to do with patient populations or it could be regulation that you know that's the part that the model doesn't tell us we need subject matter experts to help inform us on that.
34:25But it's a fascinating insight. I know, Sareb, you've got some, had some examples of that as well. Yeah, so one of the companies we're doing a bit of work with and talking to, they are doing a clinical trial around vitamin D supplements. So one of the things the model actually pulled out is if you're going to do that in Asia, you want to be in Indonesia. And probably a reasonable question to ask, because if you're in Australia, it's an Australian company. We've got obviously people from Europe came here 200 years ago. Europe didn't have a lot of sun, a lot of sun in Australia. So people in Australia tend not to have vitamin D deficiency.
35:03Plus we have a lifestyle which is very outdoorsy, very beachy, and all those kinds of things. If you're in Indonesia, it's a heavily Islamic country. So people cover up a lot and people are outside less. So they actually get a lot less sun exposure than they do here. I mean, the equivalents in America would be, look, if you're doing a vitamin D study, probably no point doing it in Florida or on a beach, but you want to do it in Alaska. And all these things are hidden in the data, which when you say it, oh my God, that's so obvious. Similar for different genetic conditions, like what condition we're looking at, the model predicted you want people from Switzerland.
Read the full transcript
35:39Because every region has slight genetic diversity of our heritage and evolution over time. And different conditions, You actually want different population sets to help you kind of really declare out efficacy or not. That's fascinating. So this is a SaaS product. Is that right? And it's customers go on to, they subscribe, I presume, and then they go on to the platform and they load their data up themselves. Is that right? Can you talk about how it works practically? So the way it works, you go on, you sign up for a trial, then if you're happy, you continue on. So the simple version is you have all those trials to search for.
36:21So you can type in a condition, look, cancer, small cell lung cancer, what are the trials currently in phase three that are going to complete in the next year? It gives you a list of all those trials and it gives you a lot of data about them. And it tells you why the prediction is 17 % or 72%. And it gives you each of those 700 variables that were meaningful, they kind of bubble to the top. But then let's say that you have your own trial. We have an interface right now. It's almost like a clinical trial simulator. We can literally upload your protocol. We'll grab whatever variables we can. You might need to tweak some, adjust them, and then you can actually simulate that trial.
36:57And what we're building now, and hopefully we'll be done in the next couple of weeks, is almost like a bit of a slider bar where we'll recommend the optimal for every variable. But you might say, look, the model says we needed 700 patients we can't afford it and you know we can only get 300 patients based on our budget so you can slide that down and we'll tell you what the impact of that is on your probability of success um for the first couple you probably need our help but over time companies will just be able to do this themselves yeah how many uh variables can you tweak around 30 or 40 that you can tweak and from you know those variables that's when we can create those the full 700 variables from those base 30 or 40 variables using natural language processing, using combinations of variables.
37:43So, but yeah, there'd be at least 30 to 40. And they vary from, you know, the simple stuff like choice of patients to the actual protocol itself, inclusion, exclusion criteria, what endpoints you're trying to meet, things like who you want to work with, like principal investigators, what sites, what countries you want to be in, all that sort of stuff is selectable and has an impact on the probability of success. So you can imagine, you know, if you're tweaking 30 variables, there's the scope to really, I don't want to use the word gamify, but let's say optimize your trial. And, you know, and often it's low hanging fruit that can be tackled.
38:24You could change, you know, from doing open label masking to double blind which is a technical thing but um just by changing that it could significantly increase the probability for a certain phase or condition um and that's what our model will shed light on how big is the market for this do you think so fingers crossed it's quite big because what i'm hoping will happen with a tool like this is i mean right now about 35 000 trials happen in the world every year on humans. I think what a model like this will actually do is make it so there'll be millions of trials that will happen in the AI world.
39:05And you'll have less trials happening in people. So it might not be 35 ,000, it might be 5 ,000. But you'll find those trials will be like overfunded and over-resourced so they can kind of get through the game very, very quickly. And at some point in time, it's going to be like, well, why would you do a trial in people if you haven't done it in AI first? it will almost become like an ethical thing like you know i'm sure kids learning to drive in in a decade like you won't learn to drive in a car you learn to drive in a simulator and once you're great at that then they'll let you out on the real road because it would be so unsafe to let someone learn to drive on the real on the real road um so i think this will just be a thing in five or ten years every trial will happen in ai and once they've optimized it and figured out which ones to back and not back that's what they'll do in the human population yeah wow that's uh That's exciting.
39:51That's fascinating. And you don't look at the molecules, though, do you? Yeah, one of the benefits of using natural language processing models is that you can extract the molecules and compounds that were used for each trial. And then there's a separate database that has the properties of those molecules and compounds. And we're in the process of ingesting that information into the model. It's one of our next steps. and that will be very interesting to see what sort of an impact, you know, those properties will have on success rates. So my hypothesis is that they will have an impact but not as much as the trial design itself.
40:37From what I've seen so far from TrialKey is about 60 % of the world's trials are designed poorly. There's major design flaws in them. So I think that there's a lot of drugs or devices that would probably have been successful had the trial had been structured correctly. It's not the underlying molecule or compound that was the problem. It was the actual design of the trial that led it to failure. So I think hidden in the data, there would be thousands of these drugs or devices that would have shown efficacy and they're just sitting there orphaned at the moment. So maybe something like TrialKey might bring back some of those prospects and candidates and better testing can be done.
41:24Yeah, wow, that's exciting. Yeah. So the product is launching when? It launched about a month ago. So you can go on to trialkey.ai, you can sign up now, you can get a bit of a trial. One other point I'll make, just further to Damon's point. So where Damon and the guys really thought, oh, there's something to this, was during COVID. So that's when they first built the model. There were about 800 vaccine candidates out in the world. And what the model predicted was Pfizer and Moderna, sorry, Damon? uh it was fight Pfizer and Moderna were I think one and three um in our predictions um and that's the way the world ended up and that's the way you know reality happened um but if you knew that before the event um it's almost like buying a lottery ticket after it's drawn then you would have just over indexed your effort and resources on on those ones and the other 750 you might have done the top 50 but other 750 you just wouldn't have done you would have really kind of pivoted your resources towards those guys.
42:28But that was an example of something really novel, right? That was all around mRNA vaccines, which hadn't really been done before, but the model still got it right. Because a big part of their success was how they designed their trials. And some of these large farmers just have more money, have more resources, have more expertise to design them better. And that's what I think will get democratised across, you know, all pharmaceutical companies. They'll all have access to the same level of insight. And particularly hospitals and universities, which are systematic underperformers when it comes to clinical trial success.
43:00The models show on that, you know, to be a fact of the data. So I think products like TrialKey could lift the standard of research in hospitals and universities and create a level playing field, I suppose, so that they could compete, you know, with the bigger end of town. uh and i would guess it would also increase investment for underfunded companies that have good ideas uh if they can show that the the likelihood of a trial succeeding uh people would be more willing to invest at some point do you think this would be standardized so that there is a trial key seal of approval or trial key badge or something.
43:53Because predicting for the company is one thing, but the company being able to use that to raise capital or whatever use they might have beyond actually going ahead with a trial would be important. Totally. I mean, I think what it will be, because the real subtlety of what you said was, I think a tool like ours would lead to an efficient allocation of capital, where money and resources will go to the trials and the drugs and the molecules that are most likely going to succeed. And once that kind of gets accepted and people see this working and see this working over time, I think you're right, it will flip the other way.
44:33So anyone who's looking to seek funding will probably have like a trial key rating. um we'd like you get energy ratings for appliances that same kind of thing there'll be probably a similar rating for all trials um and then before they go to seek funding they'll probably use a tool like ours to really optimize it so to maximize the chance of success and chance of success it's a funny thing like in that rare kidney disease 54 was amazing um but in something like obesity you really want to be above 70 to be differentiating um and then you can kind of use that optimize and it will just be a standard part of the funding process or for a lot of the large farmer to probably just a standard part of their internal due diligence as they choose which trials to call and which ones to fund i'm wondering craig whether you should join our marketing team i like the trial key seal of approval i really i'm i'm really warming to that uh yeah well that's uh That's amazing.
45:29Whose idea was this? Well, the idea came about three years ago. It was the then CEO of Opal, Michelle Gallaher, and I had a discussion around this market problem. And it was one that I couldn't believe hadn't been solved already given it's$15 million on average to invest in a phase two trial for a pharmaceutical company. I'm running sports betting models in the background and pharmaceutical companies and hospitals, universities aren't using big data for clinical trials. It just seemed like a complete, I guess, oversight that they weren't using machine learning and big data to help them design trials.
46:17So with that in mind, we set out to solve this problem and we were one of the first. there's been you know one or two competitors and some academic papers written but as i said before this is still pretty much virgin territory so um yeah coming from a sports betting background i just could not believe that this hadn't been solved or attempted to be solved using machine learning yeah and uh i'm interested uh in the investment side are you thinking about you know making selling this data to investors? Yeah, totally. I mean, we haven't quite figured out the right model, whether it's like a subscription fee you charge to investors or whether we actually just set up our own separate fund and use that to invest in.
47:04So that's why for now, I think it's just more Damon and I doing trades on the side. But once we kind of crack the pharmaceutical game, I think this will be the next use case. And we might then end up bifurcating the organizations. One's a lot more pharma focused. One's a lot more investor focused. And as Damon, as you were saying, you know, you were surprised no one's done this. This kind of a model could be applied to all kinds of things. Have you guys done any or is it too early or do you think you will? No, I've looked at other areas, other applications for this type of modeling. One I've just launched actually is called golf swings dot AI.
47:48which will, I don't know if you're a golfer, Craig, or know any golfers, but it can take your golf swing, just film from a regular smart device. It can determine all your movement points, and then use explainable AI to predict what movements you can work on to fast track your improvement, which ones are most likely to result in a handicap reduction. So we've only just launched that product, golf swings.ai. And there's other applications as well for digital marketing, all sorts of possibilities in this space. But this trial key one that we're talking about today, given the size of the industry and what's at stake, this is an absolute prime candidate for this technology.
48:33And we're excited about it. But, yeah, you're right. There are definitely lots of use cases. And, Saurabh, what you were saying about, I think, did you say 30 30 000 odd uh clinical trials a year uh and that could be reduced to 5 000 if people uh knew the likelihood of of success in a trial and then yeah you'd run them in simulation before going into human trials that's that's a pretty powerful idea so and just a bit further to Damon's point. I mean, basically what the guys have done is they've built a generic way to solve what's actually a super generic problem. You've got a thousand variables, you know, an outcome, what variables lead to that outcome?
49:22And how do you weight those possibilities? And it just happens to be applied to the clinical trial space. And that will be what we'll do for the next couple of years till we really run this problem down. But it's a generic solution to an actual super generic problem. Is there anything I haven't asked that you guys want to say? No, I think that was a really good discussion. You asked all the right questions. Hi, I wanted to jump in and give a shout out to our sponsor NetSuite by Oracle. A little quick math. The less your business spends on operations, on multiple systems, on delivering your product or service, the more margin you have and the more money you keep.
50:03But with higher expenses on materials, employees, distribution, and borrowing, everything costs more. So to reduce costs and headaches, smart businesses are graduating to NetSuite by Oracle. NetSuite is the number one cloud financial system bringing accounting, financial management, inventory, HR into one platform and one source of truth. With NetSuite, you reduce IT costs because NetSuite lives in the cloud with no hardware required, accessed from anywhere. You cut the cost of maintaining multiple systems. You improve efficiency by bringing all of your business processes into one platform, slashing manual tasks and errors.
50:53Over 37 ,000 companies have already made the move. So do the math. See how you'll profit with NetSuite. Popular demand, NetSuite has extended its one-of-a-kind flexible financing program for a few more weeks. Head to netsuite.com slash ionai. That's ionai, E-Y-E-O-N-A-I, all run together. Again, head to netsuite.com slash ionai. That's it for this episode. I want to thank Saurabh and Damon for their time. If you want to read a transcript of the conversation today, you can find one on our website, IonAI, that's E-Y-E hyphen O-N dot A-I. Take a look at opal at O-P-Y-L dot A-I. And remember, the singularity may not be near, but A-I is changing your world.
51:52So pay attention.
From the publisher
This episode is sponsored by Netsuite by Oracle, the number one cloud financial system, streamlining accounting, financial management, inventory, HR, and more.
NetSuite is offering a one-of-a-kind flexible financing program. Head to https://netsuite.com/EYEONAI to know more.
Unlock the secrets of clinical trial predictions with Saurabh Jain and Damon Rasheed from Opyl & TrialKey.ai on this episode of Eye on AI.
In today's discussion, delve into how TrialKey's pioneering AI model can predict the success of clinical trials with astonishing accuracy—about 90%. Learn about the innovative use of over 400,000 past trials to enhance the precision and efficiency of new pharmaceutical trials. Saurabh and Damon share how their technology not only predicts outcomes but also optimizes trial design, saving substantial time and resources in the drug development process.
The episode also covers the sophisticated data processing techniques employed by TrialKey, including natural language processing to extract vital variables from complex datasets, and the strategic use of these insights to improve trial designs. Discover the implications of poorly designed trials and how AI can revolutionize their structure for better success rates.
Tune in to gain a deeper understanding of how AI and machine learning are transforming the pharmaceutical landscape, making clinical trials more predictive and efficient. Saurabh and Damon's expertise offers a glimpse into the future of medical research and investment, emphasizing the role of AI in enhancing decision-making and resource allocation in healthcare.
Remember to like, subscribe, and hit the notification bell to keep up with the latest innovations and discussions reshaping the future of artificial intelligence and healthcare.
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI




