#201 Danny Halawi: The Science Behind AI Predictions - Accuracy, Challenges, and Innovations

5 Aug 2024 · 55 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. Podcast Episode #201 Summary

Episode Title

Danny Halawi

The Science Behind AI Predictions - Accuracy, Challenges, and Innovations

Host Craig S. Smith

Guest Danny Halawi, PhD student at UC Berkeley

Episode Overview In this episode, host Craig S. Smith interviews Danny Halawi about his research on leveraging large language models (LLMs) for forecasting future events. Halawi discusses the accuracy of these models, the challenges they face, and the innovations in the field of AI forecasting. The conversation covers various aspects of judgmental forecasting, prediction markets, and the potential applications of AI in decision-making processes.

Key Topics Discussed

  1. Danny Halawi's Background
  2. Transition from computer security and fraud detection to AI forecasting.
  3. Inspired by the advent of GPT-3 and its implications for society and decision-making.
  1. Automated AI Forecasting
  2. Development of an AI system to forecast geopolitical events.
  3. Implications for AI safety and aiding public decision-making.
  1. Judgmental vs. Time Series Forecasting
  2. Time Series Forecasting: Relies on abundant data and minimal distributional shifts.
  3. Judgmental Forecasting: Based on human intuition, historical data, and domain knowledge.
  4. Importance of understanding both methods for improved forecasting.
  1. Role of Prediction Markets
  2. Platforms where individuals assign probabilities to future events, like elections.
  3. Aggregated forecasts from prediction markets often outperform individual expert predictions due to the "wisdom of the crowd."
  4. Weighting of contributors' probabilities based on their historical performance.
  1. Model Development and Data Sources
  2. Use of APIs to collect real-time data from prediction markets and news articles.
  3. The necessity of summarizing information to fit within model context windows for effective predictions.
  4. Exploration of reinforcement learning for model improvement.
  1. Challenges in AI Forecasting
  2. Training models on high-stakes predictions with uncertain outcomes.
  3. Ethical considerations related to AI predictions in real-world scenarios.
  1. Future Directions and Potential Applications
  2. Growth in the field of AI forecasting and its possible societal impacts.
  3. Potential for superhuman prediction capabilities through enhancements in model training and data sourcing.
  4. Importance of continued research and investment in AI technologies to facilitate better decision-making.

Key Takeaways

  • AI forecasting can rival human forecasters and prediction markets in accuracy.
  • The architecture behind AI models utilizes extensive datasets and real-time data collection.
  • Future developments could improve prediction capabilities, particularly in uncertain scenarios.
  • Ethical considerations and practical applications must be addressed as AI becomes more integrated into decision-making processes.

Conclusion The episode concludes with a reaffirmation of the significant role AI is poised to play in shaping future decisions and understanding the world. Craig reminds listeners to stay informed and engaged with developments in AI technology.

Sponsorship This episode is sponsored by Oracle Cloud Infrastructure (OCI), emphasizing the need for robust and cost-effective solutions for AI model training.

---

For further updates and insights, listeners are encouraged to follow Craig Smith and Eye on A.I. on Twitter.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00We release a paper showing that an end to end language model system can forecast future events almost at the level of human crowd aggregates almost at the level of these prediction markets. And in some specific settings, they're actually better than the prediction market. We wait till the question resolves. And if the model made an accurate prediction, which basically means it gave a reasoning that was effective at predicting the event. Like we collect this reasoning and then we train the model on these reasonings because we found them to be effective for making correct predictions. Hi, I'm Craig Smith and this is Eye on AI.

0:38In this episode, I speak with Danny Halawi, a PhD student at UC Berkeley working on AI forecasting. Danny shares his research on using large language models to predict future events, such as election outcomes or sports results, at a level comparable to human forecasters and prediction markets. We discuss the potential of AI forecasting to help individuals and organizations make better informed decisions, the architecture of the AI system Danny and his colleagues have developed, and the challenges of improving its accuracy further. Danny also provides insights into the world of prediction markets and the wisdom of crowds.

1:23I hope you find the conversation as fascinating as I did. AI might be the most important new computer technology ever. It's storming every industry and literally billions of dollars are being invested. So buckle up. The problem is that AI needs a lot of speed and processing power. So how do you compete without costs spiraling out of control? It's time to upgrade to the next generation of the cloud, Oracle Cloud Infrastructure, or OCI. OCI is a single platform for your infrastructure, database, application, development, and AI needs. OCI has four to eight times the bandwidth of other clouds, offers one consistent price instead of variable regional pricing, and of course, nobody does data better than Oracle.

2:15So now you can train your AI models at twice the speed and less than half the cost of other clouds. If you want to do more and spend less, like Uber, 8x8 and Databricks Mosaic. Take a free test drive of OCI at oracle.com slash IonAI. That's E-Y-E-O-N-A-I, all run together. Oracle.com slash IonAI. That's oracle.com slash IonAI. Yeah, I did my undergrad at UC Berkeley in computer science. At the time, I was very interested in computer security. So, you know, after graduation, I worked at Google and Robin Hood, where I was doing machine learned fraud detection. And then, you know, around the time that GPT-3 came out, I thought it was a big turning point for AI, as a lot of people did, both from an advancement standpoint, and also just as a, you know, what are the implications for broader society?

3:21you know, as a for the labor force, just from a safety perspective. And so I felt compelled to sort of like go back to school and do work in this space. So I started working with Jacob Steinhardt at UC Berkeley. And, you know, I was really interested in understanding how models work. So, you know, we know that models work very well. We know neural networks are strong, but we don't necessarily know how they work. And so this is a problem that really gripped me because these things are deployed in high-stakes situations, but we have no idea what the underlying algorithm that they're implementing is to make decisions and do various tasks.

4:05So this is one area of research that I was very interested in. And then recently, we've done work on automated AI for automated forecasting with AI. So this is like forecasting future geopolitical events with AI systems. So, you know, we believe that this has implications for AI safety and also just implications for, you know, the broader public, how they could potentially use this technology if it were to work. So you can imagine that if we had a system that could forecast you know, future events better than humans, we would be able to understand the results of taking certain actions and making certain decisions.

4:51And if we could understand those things, we could, you know, make better informed decisions on what actions to take. So currently, you know, we have forecasters that are humans that, you know, forecast different events. So the really good ones are called super forecasters. And the thing with superforecasters is that there's not a lot of them and they're expensive. So if you can have a machine that could forecast at the level of humans, this would be very valuable because you could distribute this technology, you know, widely and people could use this to make more informed decisions. And then also you can imagine as AI gets really advanced, using AI to forecast AI itself, meaning using AI to forecast when AI will be able to do a certain thing or when it will get to this level of capability is really important so we can understand, you know, what is the world going to look like in two years?

5:54So these were the two things that got us interested and using AI for forecasting. And yeah, that's what I've been up to recently. Yeah. And so tell me about the system that you've built in the paper that you wrote. I guess start with the prediction markets, which I wasn't really aware of, and how those function, and also how the super forecasters function. Are they doing deep research around an event? Yeah, absolutely. So I guess the thing I'll start with is some context, which is the thing we're doing is referred to as judgmental forecasting versus time series forecasting. So time series forecasting typically excels when data is abundant and when it doesn't change much.

6:49So it's under minimal distributional shift. And here in this field, you take tools from time series modeling. So like you can imagine if you want to forecast the weather, you just collect a lot of data that you have currently in the environment and use some time series modeling to make a forecast about the weather. Judgmental forecasting, on the other hand, is where people assign probabilities to future events based off historical data, their intuition, their Fermi estimates and their domain knowledge. So it's like very much intuition based and it's very much just thinking about, you know, what are the possible outcomes?

7:25What are the factors that I should consider and what are the probability of the different factors? So we are interested in the field of judgmental forecasting. OK, and to study this, you can take a look at prediction markets. So prediction markets are like essentially these platforms where people put probabilities on different events happening. So anybody can go onto these prediction platforms, take a look at a question that interests them, like who's going to win the next U.S. election, and they can put a probability on each candidate. And you have many people doing this. And based off what people think, this will determine what the platform or prediction market thinks is the true probability of this event.

8:06So if you have like 500 people putting a probability on, you know, Trump or Biden or some other candidate winning and you take a look and you aggregate all this information of what all these different people think, the prediction market puts a probability on the event. So like one natural concern is like, OK, well, some humans are probably better at forecasting than others. So like, why should the platform consider everyone's probability? And the thing that these platforms do is they weight people's probabilities based off their past performance. So if you've been on these platforms for a long time and you've done very well at predicting the future, then the platform puts a lot of weight on your probability because you've done a good job in the past.

8:55So they have reason to believe you'll do a good job in the future. And another interesting thing I'll say about these prediction markets is that if you take a look at the aggregate forecasts of forecasters, so you take a look at what the prediction market says is the probability, which is based off of the individual forecasters, this probability is usually more accurate than the probability reported by the expert forecasters. And the reason why this happens is because every individual has their own biases. They base their forecast off the information that they find compelling. And when you average the probabilities across many different individuals, you're sort of getting all the different pieces of information that each individual is considering.

9:50and you're also sort of averaging out this bias factor that each individual has. So this is referred to as the wisdom of the crowd. And that's why prediction markets are a very valuable tool and are very successful. And it's because by taking what everyone thinks and averaging them in some way, you can get a more accurate forecast than you would have. So that's basically the field of judgmental forecasting. And this is sort of the state of prediction markets. I guess I'll also add that there's been, I mean, prediction markets have been around for a bit, but there's been a lot more excitement about them recently.

10:31So like just a lot more user participation. This has almost become like a trend or a fad where people go on these like prediction markets and just, you know, start, you know, making probability, start submitting probabilities for different things. and people will even like, you know, start like propose their own questions. So there's been a lot of excitement in this space and a lot of attention. And which I, which I think is a good thing for this field. You, let me just ask before we go on to, from that, the, what's the accuracy? is there sort of an aggregate accuracy for the different platforms?

11:16And I would imagine that when you break that down to different domains, they're more accurate in some domains than others. Yeah. And then if they are 70 % or above accurate, are people using them to gamble? because that seems like an obvious use case. Yeah, so I think certainly, so these prediction markets don't actually use real money. So that's important to say. You're not actually going on and placing a real bet. Although that would probably increase the accuracy of some of these questions because if you're actually using real money, maybe people are more incentivized to really think through the question.

12:02Um, people do like rely on the prediction markets for making their own decisions. Um, like people do use the wisdom of the crowd to, uh, actually like incorporate into their life when considering different things, whether it, uh, to, you know, maybe take on this major in college, like maybe, maybe like people who are entering college right now are trying to understand like, okay, if I go into this field of medicine, we'll just be completely automated by AI in two years. And so I think there are people who really like take information from these prediction markets and use it practically as far as like betting goes.

12:47So there's lots of questions on sports on these platforms. And I wouldn't be surprised if people do use them to make sports bets, but I would be surprised if anyone has the use them successfully to make money. And I think that's because when you think about the things that we can bet on, like sports events or maybe elections, these things are extremely hard to predict and very chaotic. So even if you get a more accurate probability from the crowd, which you will, there's just a level of noise that you can't really beat, which still allows basically the house to win almost every time. So, yeah, that's the current state to answer that.

13:33Yeah. But what's the reliability or the success rate on these platforms? Yeah, yeah. So up until 2022, I could give you some stats on some of these platforms. So an earlier paper by Andy Zao and colleagues looked at three platforms, Metaculous, Good Judgment Open, and CSET Fortel, which are three prediction markets. And across all questions, the crowd got, I believe, an accuracy of 92%. But if you look at the same exact platforms in the last two years, which is what we studied, their accuracy falls to 77%. This isn't to say that they're getting worse. This is just to say that the types of questions that they're forecasting are getting more and more difficult.

14:38So I would say, I mean, maybe one thing that's helpful. as a statistic is like these prediction markets, if you were to compare them, like, let's say you were to do a competition with like individual forecasters that are, you know, really, you know, have a lot of experience and are very smart and are have known to be accurate. These prediction markets on average tend to be even better than them. So the accuracy will completely just depend on the difficulty of the questions you're considering. But as far as how powerful are these markets, I would say that they're quite strong. And I think that they will only continue to get more strong as you get increased participation.

15:21Yeah. And before, again, we go into your paper and your system, just aggregating the predictions of the prediction market seems would be one step towards a stronger prediction. If you have, I don't know how many there are, five or ten markets uh probably um the the questions are phrased slightly differently and that would create a some noise but uh but yeah can you can you boost the prediction simply by aggregating the prediction markets themselves yeah absolutely i mean sort of like one of the key takeaways from the forecasting literature is like the more sources of information you have, the more signals you have, the better your forecast is.

16:13So there are actually currently platforms. I can't remember the name of the one I'm thinking of, but there is a current platform that simply aggregates the prediction from all the prediction markets for a particular question. So this is, yeah, this is done in practice. And I think one thing that you pointed out is you have to be careful because the questions might have different wordings on different platforms, which would create different probabilities for the event. Or they might have different resolution criteria. So like one platform, like two platforms will have the same question, but maybe one platform will say like, this will resolve as yes, under this criteria.

16:53And then another platform will say this will resolve as yes, under this criteria. For example, let's say the question is, will the Starship launch be successful. You know, people have different criteria for what successful is in one platform might outline, you know, certain criteria for a successful launch and the other one will have different, different criteria. And so, uh, yeah, you just have to be careful with respect to that. Um, but overall this, this would be a very intuitive and natural thing to do. Yeah. Uh, okay. And then, uh, let's talk about what you have been doing with, uh, large, large foundation models.

17:30Cool. Yeah. So recently our group at UC Berkeley, with Fred Zane, Yuhan Chen, and Jacob Steinhardt, we released a paper showing that an end-to-end language model system can forecast future events almost at the level of human crowd aggregates, almost at the level of these prediction markets. and in some specific settings, they're actually better than the prediction markets. So yeah, I could tell you about, we could start with the data set collection and I could tell you about what the task is and how, and the motivation for it and why we set it up this way. And then I could tell you how our system performs.

18:15So on the data collection front, there's all these prediction markets in the wild and most of them have APIs or it's easy to you know scrape the data online and they allow you to do this and you could scrape the questions you can scrape what the crowd thinks at different dates so for you know the crowd prediction is constantly changing over time as a new information is accrued it's not like there's a question they asked the crowd once and they use that as a probability like they asked the crowd at, you know, across a question's open date and close date. So the question might be open for six months or something.

18:53And the crowd prediction changes. So you can get all this data. And then given that you have these questions and this data, you can ask an AI system like a language model what it thinks the probability of this event happening. And you can compare it to what the crowd thinks. And then you could just see, I mean, eventually at some point the question is resolved, the outcome happens, and you can sort of compute the accuracy across many different questions over time and see how accurate your system is. So this is the setup. So we thought that language models would be like a really good place to start with getting, you know, superhuman machine forecasting.

19:37We thought language models was a good place to start for a couple different reasons. The first is that, like, for human forecasting, typically humans or expert forecasters or people who do this as a profession, you know, forecast in specific domains. So if you want really good forecasts about the environment, like, there are certain people who are really good at this. But this limits you because you have to rely on experts. And some questions, you know, call upon interdisciplinary thinking, which isn't always trivial to think about. So language models are trained on, you know, web scale data. So they're endowed with mass cross domain knowledge.

20:18So they're not, they don't, they don't succumb to this limitation. Furthermore, they can parse and produce texts extremely rapidly. So they're timely, you know, you can invoke them at any point in time, you know, during the day or at any point in time during, during your decision process. And lastly, they can be trained in a way that humans can't. So language models as they currently stand are trained on some very large dataset and then deployed and then used. But once they're deployed, they're not getting updated or retrained with recent knowledge. I think good evidence for this is if you go to ChatGPT, You ask it like who won, you know, the last NBA finals, I won't be able to tell you because it was trained up until 2021.

21:10Once it was released, it wasn't trained again. So why is this useful and interesting? It is because I can ask a language model about events that have already been resolved or have already happened and see what its prediction is and compare that to the crowd prediction. OK, in addition, if I let it look at information that was only available up into some date, I can sort of simulate forecasting. So like today, if I were to forecast the state of the election, like who what I who who I think is in the US election, you know, go online and I have information available up until today, March 8th. Right.

21:53And tomorrow I have information available up until March 9th. And so you could, you know, retroactively go to these, you go to these models and just give them information up into some date and then have them make a prediction. And this will effectively be like, okay, what did, like, what would this for, what would this model have said if it was really forecasting on this date? So you can set up this method of evaluation and method to train these models on information that has already been resolved in a way that's automatic and effective. So, like, one method you could even consider is having a language model that was deployed and trained up until, you know, with data up until 2023, to sort of teach and guide a model that was trained only up until 2021.

22:43The point here is that the 2023 model would have so much information about things that have happened that it could, you know, have the earlier language model make predictions and it could help you, you know, correct its predictions and learn from those learn like it could help it learn where its mistakes are and how it could have made a better prediction and improve accordingly. So you can have this, you know, completely automatic training between two systems where there's literally no human involved. So we thought that this would make language models like a really good place to start for forecasting.

23:18Interestingly, like this idea has been proposed and tried in the past. So like we're not the first people to like come up with this idea or try it. And, you know, in a study that was released up in 2022 by Andy Zao and his collaborators showed that, you know, their system achieved an accuracy around 65%, whereas the crowd accuracy was 92%. So their system was just like a lot worse than the crowd. But as language models have become more capable over the past two years, we thought to try this again, potentially with better, you know, sourcing even better information and, you know, training this system in more advanced ways.

24:08And what we observe is that, like, the model has gotten more accurate. So our accuracy was 71 % versus the crowd accuracy of 77%. And what we hope to show here is that language models are a good place to train, you know, machines to forecast. And we believe like Like this gap will only become more narrow over time and eventually machines will become better. And if this is the case, we hope to make, we hope that this technology is distributed across organizations and people so that they can use it to make more informed decisions. Yeah. Yeah. And the data. So let's say you want to predict, you know, a sports team, not that I'm particularly interested in gambling.

24:57But you have the Super Bowl. And are you relying on the information encoded in the weights of the models? Or are you supplying the models with reams of data on individual players, on historical outcomes? on what kind of data are you providing and how do you collect that data? Yeah, great question. So I'll answer that question. Then I'll also talk a little bit about how our system was trained since that'll give also insight on what sort of data the model is privy to. So yeah, so at test time, when we're actually having the model make forecasts, what we do is we give it access to essentially the internet.

25:52so up until some date, of course, because we don't want it to have information after a date. And what it has access to is what we let the system do is decide what it wants to search on the internet. So for forecasting, let's say the Super Bowl, the system will say, okay, I want to see what experts are saying. I want to see who's currently injured. I want to see what the weather is going to be like for the game. So the system itself will come up with everything that it wants, what types of information it'll want to see. And then we go and grab that information. And then the system looks through, let's say, 300 articles, and it determines which ones are relevant and which ones are irrelevant.

26:32So when you go try to look up the answer to a question on Google, you might have to look through a few articles before you actually get the answer to your question. So the system does the same thing. We grab a bunch of articles and we decide which ones are like most relevant to the questions it was asking. Okay. Let me just stop. When you say grab the articles, are you talking about physically downloading content and then uploading it to the model? Are you talking about copying URLs into the model? Or yeah, how does that happen? Yeah, so we grab the content of the article. And the way we do this is we use news APIs.

27:11So there's a bunch of news APIs out there. Google News is one we use. Another one, there's a startup called News Catcher, which we think sources like excellent information. I expect that to do very well. And they provide, you know, the body content of an article. So, you know, we load this into a data structure, all the content in each individual article. And then we have the model go through each article and, you know, look through the article end to end and see if the information is relevant. So, like, this is step one. Okay. And let me stop you there. When you say you, are you putting the content into, like, a vector database or are you?

27:57Yeah. Uploading PDFs to the model or how? Yes, we're feeding it into its context. So just like you prompt it, like just how you go on like some, you know, web app and prompt a language model a way that we typically do as, you know, just common users. This is exactly how it's being done. So we're not like embedding the model. We're not embedding the article in a special way or encoding it in a special way. We're just giving the plain text to the model within its context and then asking it some questions like, do you think this is relevant? given like what you wanted to know. Does that make sense?

28:34Yeah. Cool. Um, so yeah, so once, so once you determine which articles, once the, once the system determines which articles it finds relevant, the next step is to summarize this information. Why do we summarize this information? It's kind of for the reason you just said. So we're giving this information to the model within its context window, um, which might be 16 ,000 tokens or 32 ,000 tokens. So we have a limited space to work with. So we take like the 30 most relevant articles and we summarize them to condense the information. And then we give it to like a final language model and say, okay, look, these are all the retrieved articles based off the questions that you ask, based off what you found was relevant.

29:15Finally, like apply these like forecasting methods to make a forecast. These forecasting methods could be like, what do you think the reliability of each information is. What do you think, like for each piece of information, how much weight do you put on this for affecting the outcome? And like, what is the historical base rate of this event happening? You know, that's another thing it might consider. So it just applies all these tools from forecasting and then makes a prediction. So that's how the system works in deployment. I've been talking to journalist researcher and some data scientists about doing something similar, gathering information around a news event and seeing if a large language model or an ensemble of models could determine the probability of truth of any particular narrative around that event.

30:24And the point is that there are narratives that are not adopted by the dominant narrative. There are sub-narratives that are excluded. And, you know, you can think of a million examples in a conflict. There's, you know, eyewitness accounts on Twitter or social media that don't necessarily get picked up because they're considered too anecdotal or too biased or whatever. And in the same way in this prediction, as you add data, presumably the prediction will change. I mean, if you're using the news sources, that's one source of information. But if you use something noisier like Twitter or, I don't know, I'm not young enough to be using social media that much, but whatever platform is active on a question, you may get a different prediction.

Read the full transcript

31:43So have you guys taken that into account? I think thinking about what sources of information tend to provide more reliable data is very important for forecasting. We let the language model sort of learn on its own what information it finds to be effective to rely on. One interesting thing is, I mean, one way you can measure the truthfulness of data is how well it can be used to predict future events, for example. Yeah. So if you find that, like all the information you source from a certain news outlet is always like really helpful for improving the accuracy of your predictions, like this is sort of saying, OK, well, this information, I mean, this is a good proxy for the truth of information.

32:29Yeah. One follow up that I've gotten from people along the same lines that you're that what you just said is like using this for epistemic. So, I mean, one thing you could ask is, okay, given this language model, if I can source information from all these articles and sort of like reason about this information, how many people's minds can be changed using this language model? If I give this language model to like a thousand people asking them, you know, some controversial question and the language model can pull information from different news sources. You know, one question you could ask is like, how many how many people's minds can it change on a certain topic?

33:09So, yeah, and I think measuring as far as like how this is measuring the truthfulness of information, I think forecasting is a good starting point for this, because if you rely on information that's always leading you to make like inaccurate forecasts of future events, you could sort of determine that it's not very accurate. And during training of our system, like you could probably hypothesize that the model did learn which information to sort of rely, like which sources to rely more on and which sources not to rely more on. And you don't have to like train explicitly to do this. It will implicitly learn to do this because it's trained to make accurate predictions and better information will give you accurate predictions.

33:55Yeah. So you're relying solely on generative models or are you, I mean, this sounds like a perfect use case for reinforcement learning. So I think reinforcement learning has a lot of potential applications here to make this model stronger. Currently, we're just working. We're using autoregressive language models. So generative models. And we're having the models rely on their pre-trained knowledge. Anything it knows about, you know, the context or the specifics about the thing we're asking about, as well as pulled knowledge. But you could, for example, let's say you had access to a lot of humans.

34:40You could have the model generate four forecasts for an event. It'll just generate four different types of forecasts. And you can have a human who's maybe good at forecasting rate each forecast across different criteria. Or you can have even the model do this itself. and then you could train and then you could use some form of reinforcement learning to like reinforce the criteria that was like most effective so yeah there's we think that our work is preliminary and we imagine that there's going to be a lot of work done in this space after sort of like the release of our work I think people were not very bullish on like forecasting is hard like predicting the future is a very hard thing to do And so I don't think folks have been bullish on AI being able to do this.

35:29But I think given the advancements and training algorithms, which you alluded to, and just the capabilities of the size of these models, I think this will become very possible in the near future. Yeah. So your system, and can you just at a 40 ,000 foot level, give us the architecture? Yeah, yeah, definitely. So I'll tell you the architecture at a high level, and I'll tell you at a high level how it was trained since I think it's interesting. so I sort of was talking about this earlier but the first step in which is really important in judgmental forecasting is like what are the different things you would need to consider so if you're forecasting an election okay it's one thing to just look at what the polls say but there's a lot of different factors that go into it that go that make an accurate prediction Like maybe you want to look at what the campaign finances are.

36:31Like you want to look at recent scandals. You want to look at the state of the economy. Like what are different geopolitical events going around the world? And what are these candidates thoughts on it? So the first step is to get a model to be able to really think about which considerations are important. Once this is possible, and this is all generative. So you give the question and like information about the question to a model and you say, okay, like generate the sub questions that you have. And then you take these sub questions, and you just go to some news API, like Google News or NewsCatcher, which I said, and you look up articles that hopefully have information about these sub questions.

37:10And then once you get articles, there's going to be some articles that are just not relevant, just based off like noise in these, you know, APIs, and the fact that these search engines are not perfect. So you really want to get the best information. And then you want to condense it into the context window of a model. And this is simply, this has to be done because models have a fixed context window that you can't see past certainly. Yeah. And just up to that point, is this all automated? So you, you, you put in the question, the model, uh, makes an API call and, and gets in at all comes in. Yeah.

37:49Okay. Yeah. So the only thing we do, so it's completely automated. It's automated end to end. It takes a question and every step that I just described is a step that it does completely on its own. And then it gets this large list of summarized articles and, you know, which with each article, like potentially having information about each sub question. I think like this on its own is a pretty interesting tool. Like you can imagine if a journalist is doing like you know research on some topic or someone's doing a literature review about some research thing you know research thing uh like having this is pretty helpful so like just as a sidestep like there's other organizations like the forecasting research institute which have done some studies on using language models to actually help human forecasters and they've shown that like by giving them access to these language models like forecasters are much more accurate.

38:46So like they don't even have to be used in isolation, I guess is my point. To go back to the architecture, you know, once you have all the information, which is, you know, obviously, you know, most the piece of the forecasting problem, you have to make considerations about the information and then finally a prediction. So we, so in our, in our paper, we focused on training a language model to make accurate predictions given information. So the way we did this is you take like a lot of questions, a lot of forecasting questions from these platforms, and you have the language model make predictions for each one of these questions.

39:30And for each prediction, it's going to give you a reasoning why basically the reasoning will sort of describe its step-by-step thinking. And then this reasoning will lead to some final probability. So it's going to be like, let's say, a half a page of text where it gives like sort of a description of what it's thinking about. And then at the end of the text, it's going to have a number indicating what it thinks the probability is the event. So you do this on a lot of questions. And then you take the questions where the model predicted the event accurately. And then you go back to the system and you train it on these prediction because these are the accurate ones.

40:09So you basically go back to it and you say like, these were the accurate ones. I want you to learn to do, I want you to learn to do this more often. And you could just keep repeating this process. So like if you had access to a lot of compute, which we didn't, and a lot of data, you can imagine just doing this like repeatedly and getting something that's very, very strong. Okay. And then, and, and, and so it spits out at the end a probability. And then you go back and, I mean, if you're doing this with a cutoff date of, you know, February 29th, and the event resolves on the 1st of March, you could run the prediction on the models or the model based on information up until February 29th, have a prediction, and then see on the first whether or not that prediction was accurate.

41:17What is that? You were saying earlier, I think 71 % compared to the prediction market, 77%. I mean, where have you gotten with the level of compute that you have available? Yeah, so what you described is, I think, a really good summary of our process. Like, we wait till the question resolves. And if the model made an accurate prediction, which basically means it gave a reasoning that was effective at predicting the event. like we collect this reasoning and then we train the model on these reasonings because we found them to be you know effective for making correct predictions um so yeah and then as far as like where our performance stands so uh one thing i will say about our system that i haven't said yet is that it's very good at forecasting events that are more uncertain so we found that it's actually better than the prediction market.

42:21So if you take questions, okay, where the prediction markets put the probability at between 0.3 and 0.7, so somewhere in that interval, you know, it could be 0.4, it could be 0.65, but somewhere in the interval 0.3 and 0.7, if you just look at those questions, our system forecasts them at a level better than the crowd aggregate. um the reason why this is probably happening is because our uh so the the model that we trained which is gpt4 uh was trained using rlhf which basically um made it so that it doesn't output predictions that are extremely confident so the model is hedging even when the evidence is extremely clear.

43:10So like, let's say the prop, let's say the crowd puts the probability of it raining tomorrow, like 95%. I don't know. But the model won't want to tell you 95 % or 96%, because it doesn't want to give you probabilities that are extremely certain, because it doesn't want to be wrong, and then, you know, be punished punished for that or something. So but for these events where the probability is like more uncertain between point three and point seven, You just look at these questions, our accuracy is better than the crowd accuracy. And then on all the questions, yeah, so we got 71%. The overall performance is 77%.

43:46That gap is certainly smaller than the previous gap of 65 % versus 92%. But I think there's still some ways to go. I think getting to the point that we did is easier than getting to the next point, which is at the level of human forecasters or even better. I think this is going to be increasingly challenging and is going to require better information, more compute and potentially better models. Yeah. And just on that, on, on, on doing better on highly uncertain questions that could also just reflect how poor the human forecasters are on those questions. Right. That's true. That's true. I mean, it's not that the model is so great at making the – because I don't know what the percentage accuracy you would have on that kind of a question.

44:45But if it was the crowd was between – what did you say? 0.3 and 0.7. 0.3 and 0.7 and the model beats the crowd and it's 0.4, you're still a long way from having a useful prediction. That is true. I think that's a fair point. I would say that my intuition is that it's not necessarily the case that the crowd is not as good at forecasting these questions. I think these are just questions with higher uncertainty. So like if you're trying to predict the last Super Bowl, it's going to be somewhere within that interval. Like you're not going to say like 49ers 95%. This would be a bad prediction. But it is to say that, you know, on that, like with these marginal differences in probability, our system is reporting something that is a bit more accurate on average since we're considering, you know, over a thousand questions and over 3000 forecasts, etc.

45:46So it's not purely just noise, I would say. Yeah. Let me ask you something about just standard supervised learning, which now sounds like old-fashioned AI. A couple of years ago, I got to know a no-code company called Accio, who I'm a great fan of. and they you can load in your your data at that time and now they they work with less structured data but you know your tabular data and you and I used it as an experiment for an article I wrote on horse racing. I mean, what they were focused on was sorting lead generation leads for sales teams. So you have 10 ,000 leads. You don't know where to spend your time.

46:58You can't call 10 ,000 people. So you train the model based on past leads and as many columns of attributes or I guess you go on parameters as you can find on each lead for which you know the ultimate outcome. and then the model learns to predict relatively accurately which leads are likely to convert to a sales. And then when you get a bunch of new leads in, you put it through the system, it ranks them according to their probability. So I was using horse racing data, of which there is a lot. I can't remember how many columns we had, but we loaded it all in. And the last column is, first of all, we trained it on past races.

47:58Yeah. And then we loaded all these columns in. And it was a classification algorithm. You know, is this going to win or not win this horse? and then you would rank the probabilities. And it did an excellent job of predicting race time odds. And we would do this day before. And, you know, that was pretty amazing. Every now and then it would predict a win for a long shot. And I was betting, you know, I think it was$2 bets on the top. I can't remember how I sorted it, but to see how I did. And, you know, I did pretty well, except that if you're betting on the favorites, you know, the favorites aren't always going to win.

49:04So over time, you lose money. But if you bet only on the long shots because they pay so much more, you don't have to hit many of those correctly to make money. And then I lost my data source. I've always wanted to go back to it. But something like that could be another data point for your system, right? I definitely I think what you're alluding to is that there is so many different places in life where there's prediction involved. And this means that there is so much data, like just let's just take horse racing. Like this could have been part of the thing we trained our model to do. And you can imagine that if you get really good at predicting horse racing, like that probably gives you some skill on predicting other sports events.

49:53So, yeah, like I think another use case that's sort of similar to what you just said is like, imagine you take every company that ever tried, you know, to launch some product, and you take the ones that failed and the one that succeeded, and you train the model to classify whether it's going to succeed or fail based on like every single company that has ever tried to like go public or do something. and you could try asking the model about a new idea with your business plan and seeing what it thinks the probability of success is. And this is just another way to teach the model how to forecast this type of event, which I imagine would allow it to gain skills that are useful for other types of forecasting.

50:31And this is why I think that using AI and language models in particular to learn how to forecast is we're in a current position to do this in a way that we haven't been able to previously with the amount of data and compute available. And I think could really revolutionize like, you know, technology and how we humans even make decisions and like assess, assess things across the board. Yeah. So, uh, so where do you go from here? You're over 70 % against the forecasting platforms and super forecasters, is that right? Right. What are you going to do with this? This is a research project. Are you taking it further, marshalling more compute, more data, adding reinforcement learning or supervised learning classifier as another data point?

51:31I mean, what's the next step for you guys? Yeah, so my prediction is that some large company, like maybe OpenAI, is... I'm actually already familiar with a large company that has contacted me that is really interested in this work and want to put a lot of resources into it. So I think there is going to be at least one large company your organization that tries to push this to the superhuman level um and then yeah like i think my main hope is that this technology happens i think it has a positive consequence for the future um and if if i didn't think that other people were going to work on this uh i would probably do it myself but um but yeah that that's that's where i think this project is headed yeah Yeah.

52:26And in how these predictions could be used, if you had a high degree of accuracy in predictions, you could modify behavior. Precisely. Yeah. And there could be sort of a self-reinforcing loop there. Exactly. I think there's some equilibrium that you'll find once everyone has perfect information. That's it for this episode. I want to thank Danny for his time. If you want to read a transcript of today's conversation, you can find one on our website, IonAI. That's E-Y-E hyphen O-N dot A-I. In the meantime, remember, the singularity may not be near, but A-I is changing our world, so pay attention.

53:18AI might be the most important new computer technology ever. It's storming every industry and literally billions of dollars are being invested. So buckle up. The problem is that AI needs a lot of speed and processing power. So how do you compete without costs spiraling out of control? It's time to upgrade to the next generation of the cloud, Oracle Cloud Infrastructure, or OCI. OCI is a single platform for your infrastructure, database, application development, and AI needs. OCI has four to eight times the bandwidth of other clouds, offers one consistent price instead of variable regional pricing, and of course, nobody does data better than Oracle.

54:06So now you can train your AI models at twice the speed and less than half the cost of other clouds. If you want to do more and spend less, like Uber, 8x8, and Databricks Mosaic, take a free test drive of OCI at oracle.com slash IonAI. That's E-Y-E-O-N-A-I, all run together. Oracle.com slash IonAI. That's oracle.com slash IonAI.

From the publisher

In this episode of the Eye on AI podcast, we dive into the world of AI forecasting with Danny Halawi, a PhD student at UC Berkeley.

 

Danny shares his groundbreaking research on using large language models (LLMs) to predict future events with accuracy, rivaling human forecasters and prediction markets.

 

Danny recounts his journey from studying computer security and fraud detection to exploring the potential of AI in forecasting geopolitical events and beyond. He introduces us to the sophisticated architecture of his AI system, which leverages real-time data from prediction markets and advanced machine learning techniques to generate reliable forecasts.

 

We explore the complexities of judgmental forecasting versus time series forecasting and how AI can enhance decision-making in various fields. Danny discusses the challenges of training AI models for high-stakes predictions, the role of super forecasters, and the fascinating dynamics of prediction markets. He also sheds light on the ethical considerations and future possibilities of integrating AI into our decision-making processes.

 

Join us as we delve into the future of AI forecasting, the potential of superhuman predictions, and the exciting developments that could reshape our understanding of the future. Don't forget to like, subscribe, and hit the notification bell for more expert insights into the latest AI innovations.

 

 

This episode is sponsored by Oracle. AI is revolutionizing industries, but needs power without breaking the bank. Enter Oracle Cloud Infrastructure (OCI): the one-stop platform for all your AI needs, with 4-8x the bandwidth of other clouds. Train AI models faster and at half the cost. Be ahead like Uber and Cohere.

 

If you want to do more and spend less like Uber, 8x8, and Databricks Mosaic - take a free test drive of OCI at https://oracle.com/eyeonai



Stay Updated:

Craig Smith Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI



(00:00) Preview and Introduction

(01:28) The Importance of AI

(02:50) Danny's Background and Interest in AI

(04:01) Automated AI Forecasting and Safety Implications

(07:34) Judgmental Forecasting Explained

(11:01) Accuracy and Challenges in Prediction Markets

(16:01) Aggregating Predictions for Better Accuracy

(19:25) Data Collection and Model Accuracy

(23:18) Improving Model Accuracy Over Time

(25:31) Data Sources and Model Training

(29:20) Summarizing Information for Predictions

(34:08) Potential of Reinforcement Learning in Forecasting

(37:50) Automating Information Collection and Summarization

(39:01) Training the Model for Accurate Predictions

(45:04) Challenges with Uncertain Predictions

(50:14) Potential Applications and Future Directions

(52:26) The Future of AI Forecasting and Its Impact

More from Eye On A.I.

All 266 episodes
#201 Danny Halawi: The Science Behind AI Predictions - Accuracy, Challenges, and InnovationsEye On A.I. · 55 min
Listen in VO