LLMs for Equities Feature Forecasting at Two Sigma with Ben Wellington - #736

17 Jun 2025 · 1 h

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The TWIML AI Podcast Episode 736: LLMs for Equities Feature Forecasting at Two Sigma with Ben Wellington

Episode Overview In this episode of The TWIML AI Podcast, host Sam Charrington interviews Ben Wellington, the Deputy Head of Feature Forecasting at Two Sigma. They discuss Two Sigma's end-to-end approach to leveraging AI for equities feature forecasting, how they create features, manage data, and make predictive models for trading and investment purposes. Key topics include the use of multimodal LLMs, strict data timestamping, and considerations for build vs. buy decisions in AI technology.

Key Concepts and Discussions

Introduction to Ben Wellington

  • Background: Ben has a background in Natural Language Processing (NLP) and machine translation, having joined the field before it became mainstream.
  • Current Role: He oversees feature forecasting at Two Sigma, focusing on predicting future prices of assets through data features.

Feature Forecasting at Two Sigma

  • Definition: Feature forecasting involves identifying and quantifying features (observable facts) related to stocks, currencies, etc., to predict future market behavior.
  • Example: Monitoring parking lot traffic at retail stores via satellite imagery to forecast sales.

The Process of Feature Creation

  • Inquisitive Approach: The team uses a detective-like approach to identify potential features from various data sources, including job postings, and tests hypotheses using historical data.
  • Importance of Historical Data: Emphasizes the necessity of collecting historical data for meaningful predictions and avoiding temporal leakage.

Temporal Leakage and Data Timestamping

  • Strict Timestamping: Critical to ensure that data used for forecasting does not include future information that could skew results.
  • Real-world example: Historical data from vendors may not be trustworthy if it includes future-known events.

Role of AI and LLMs

  • Impact of LLMs: LLMs and Generative AI significantly accelerate the process of feature extraction and hypothesis testing.
  • Current Trends: Transitioning from traditional methods to utilizing LLMs allows for rapid experimentation and iteration on feature ideas.

Innovation and Build vs. Buy Decisions

  • Open-source Preference: Two Sigma prefers open-source models for control over data and model behavior, avoiding biases introduced by pre-trained models.
  • Continuous Evolution: The fast-paced nature of AI development necessitates a flexible and adaptable platform to accommodate new technologies.

Challenges of Financial Modeling

  • Noise in Financial Data: The inherent noise makes predictions difficult, and models can easily overfit if not properly validated.
  • Interpreting Predictive Models: While interpretability is essential, some complex models may not be fully explainable, yet they can still provide valuable predictions.

Future Predictions

  • Emerging Trends: Excitement about the rising efficiency of LLMs and their potential to automate and enhance research and predictions.
  • Agentic AI: Exploration of AI systems that autonomously interact and leverage specialized knowledge from different domains to improve forecasting.

Conclusion

  • Ben Wellington shares insights on the future of AI in finance, noting the balance between leveraging advanced technologies and maintaining ethical and operational standards. The conversation underscores the exciting potential of AI while recognizing the challenges it brings to workforce dynamics and the finance industry.

Takeaways

  • The increased efficiency offered by LLMs enables faster and more innovative approaches to feature forecasting.
  • Strict data management and timestamping are critical to ensuring the reliability of financial predictions.
  • Open-source models provide flexibility and transparency, which are essential in a competitive financial landscape.
  • The rapid evolution of AI requires organizations to remain adaptable to leverage new technologies effectively.

For more details, visit the [complete show notes](https://twimlai.com/go/736).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Ultimately, our goal is to find data and features and predictions in as many places we can. right essentially collecting data historically then what i can do is i can go back in time and for every day i can see what the world looked like on that day and i can use it to see what happened to that company over the next week month year whatever maybe i'm going to end up using that as a trading signal eventually and build what we call a model which is basically something that takes data in and makes predictions about future prices of the assets we trade

0:43All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Ben Wellington. Ben is deputy head of feature forecasting at Two Sigma. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Ben, welcome to the podcast. Thanks. It's nice to be here. It's great to have you on the show. I'm looking forward to digging into our conversation. We're going to be chatting a little bit about how LLMs and Gen AI are changing the way you approach problem solving there. To get us started, I'd love to have you share a little bit about your background.

1:19Sure. Yeah. My background is in natural language processing. So, you know, that's anything having to do with computers interacting with human languages. And I like to say I got into it before it was cool, before Siri, before, you know, it just sounded neat when I read a description on NYU's website of a research area. So I didn't realize that I was stepping into this large part of the future. But but hence, that's what I did. Here we are. Yeah, here we are on this podcast. Right. I mean, who would it get? So I did my PhD in machine translation. And it was a different time back then in this space.

2:02You know, looking back, there was this tension between empirical natural language processing, which is like, if we have enough data, we can solve all of our problems. And then sort of the more syntactic approach. No, no, no. We need to understand noun phrases and verb phrases. And, you know, even so much in the 90s, right? UPenn hired a bunch of grad students to draw sentence tree diagrams to make the Penn Treebank, which was like the preeminent data set. I mean, if you were in natural language processing in that period, you knew the Penn Treebank, right? That was, you know, and you slept thinking of verb phrases and adjectives and, you know, you had nightmares.

2:38And so it was, yeah, it was a cool time. And there was this tension, how much do we really need to know about language to sound like we're speaking language, right? Do we need to, do we understand the depth of, of syntax, or can we just go the data route? And I think looking back the data route clearly won out and we can talk about that journey today, but it's definitely been, it's been interesting. Tell me a little bit about your journey to and at 2Sigma. Your title is deputy head of future, sorry, I keep saying future forecasting, feature forecasting. What is feature forecasting there. So at Two Sigma, we're an investment manager and ultimately we're trying to invest in, in instruments that we trade that, and, and to do that, we want to predict the future price of things.

3:31Right. So to do that, we have to kind of, whether it's a company or a currency, right. We have to understand a lot of information about say a stock, let's say, you know, IBM and, uh, and use all that information to make predictions about the value of that company. And to do that, we think of breaking down the world into features, where a feature is just sort of a interesting, observable fact of interest about something. We have philosophical debates about what is a feature. And so the idea is that for a company like IBM, there are thousands, tens of thousands, maybe millions of features that might tell you about it, right?

4:19It's sales numbers, the number of computers, the, you know, the people who work there, where they came from, the board, the stock market prices. And for each of those, if we can quantify them, if we can measure them and record them, then we can use them to predict what happens next. And so this idea of feature forecasting is to make a lot of interesting features about the world and then use those features to forecast about what's going to happen in the world. So this reminds me of the, I don't remember when we started doing this, but, you know, we, the royal we, I guess, but like the, we're going to take satellite imagery as it passes over Walmart and count the number of cars in a parking lot and use that to predict share price, that kind of thing.

5:03That would be an example of a feature, which is like the number of cars in the parking lot at any given hour would be a feature for sure, right? That's a very interesting feature you could get from a satellite image. That's a great example. And it sounds like it's maybe correlated, you could say, with the way we use feature for in the context of a machine learning model, but it's not necessarily like it's meant it's more like an abstraction, an abstract idea. But they end up being machine learning model features like it's highly correlated. I mean, we think the term probably derives from our use of machine learning on all of this data.

5:45So it's not just core. I mean, it's probably accurate to say that's the connection that ultimately you can do a lot of things. I mean, I can take that satellite example you just said and do something very simple with it, which is to say, all right, well, I see 30 % increase in cars in the parking lot. So I'm going to predict 30 % increase in sales. You know, that's not machine learning. It might be if you use machine learning to extract the car count from the, to make your feature, right? But ultimately that feature is going into some linear math that's quite simple, but it's still a feature, right?

6:20And another way you would do it would be to combine that with thousands of other pieces of information into a neural network. And then it's truly a feature in the sense that you're referring to. So yeah, I mean, there's a lot of overlap. I want to kind of understand like, you know, how you, maybe even like a day in the life of like a feature forecaster, like how do you, how do you go about the role? How do you think about the, you know, prototypical, you know, workflow that you're working on? Ultimately, our goal is to find data and features and predictions in as many places as we can, right so what does the day in the life look like i don't know you're walking by a target and you see a help wanted sign in the window and and then you say oh i don't know they're hiring here that they're interesting that's growing and then you say oh well you know maybe that means they're doing well and then you say oh well can i measure employee hiring across all companies that we trade well how would i do that okay and you do a bunch of sort of and then and then and then following this sort of inquisitive detective style work until you might say, okay, hey, look, companies post job postings on their websites.

7:31So I can start to collect that data. And if I can analyze it, I can start to create features about open roles for every company. Once I do that, it's not that complicated to think about, all right, I have all this data. What are some hypotheses that might come out of hiring? Is there an increase in hiring seasonally since last year? Do we see this movement. And what we tend to do is start with a hypothesis. We test that hypothesis by essentially collecting data historically. So not only do I need that job, this is the hard part, not only do I need the job postings there today, but in theory, I need historical views of those job postings.

8:13And so let's pretend for a second that I have 20 years of companies job postings, which is hard to build, but let's just give me that. Then what I can do is I can go back in time. And for every day, I can see what the world looked like on that day. And I can use it to see what happened to that company over the next week, month, year, whatever. And so I can test my hypothesis. Hey, is it true that an increase in job postings leads to an increase in share price? And it's just a scientific statistical test. If the answer is yes, maybe I'm going to end up using that as a trading signal eventually and build what we call a model, which is basically something that takes data in and makes predictions about future prices of the assets we trade.

8:53And that 20 year, I can't stress enough, where do you find 20 years of history on a company's website? That is the interesting, challenging part about this work is timestamps are so crucial. We typically, in order to get 20 years of data, a lot of the time you have to say, okay, let me start recording this and I'll come back to it 20 years later, which is a very forward-taking way. In three years, you'll be able to do your job. Right, right, exactly. So, you know, obviously 20 years is a big, but maybe there are many things that I've been recording for five, six, seven, eight, 10 years that we're now able to do, or even 15, right?

9:32So you have to have, people used to ask me, hey, what data should you be recording? What should we record? And the answer was always, anything that might be interesting in the future. I don't care if it's interesting now. It's like, this is a time capsule. If you hit record today, hey, this is a gift to either you in the future or another researcher at the firm who's gonna suddenly have this rich history of data to then go and study. Because if you don't have that historical data, a lot of data just disappears like forever. There's no recording of it. There's a great example, which is we worked with a news, you know, a newswire, and they would send us news stories about companies.

10:10And one day we get this call, hey, I heard you have a recording of our news. Can we buy that from you? Because a news company, they don't mind rewriting their headlines. They don't mind changing, adding little topics, editing. And so they don't actually know at noon on this day what they actually had sent. They only knew the edited version. And so we kind of had this beautiful real-time history of the world that many people didn't start recording. You know, in the early days of the company, people didn't really realize how important recorded history was. I've got to imagine that the hard part of that, what should we record or what should we save is really what shouldn't we save?

10:54That's true. I think today there is what shouldn't we record. And obviously there's legal issues that you have to clearance to record something. And so there are things that would prevent you from recording for various reasons. but in terms of things that are open to the public and available um yeah if you think it might be interesting maybe kind of sort of like let's record it and then figure it out later don't worry just make sure you have access to the rawest data possible because the more you analyze it and throw out the raw data the less useful it will be in the future if someone wants to do something different you know you talked about the walking down the street seeing target you know uh sign that's that's very organic scenario but i've got to imagine at the scale you're operating at, you're being approached constantly by companies trying to sell you data that they've already collected.

11:42Like how much of your job is trying to figure out these organic sources versus, you know, parsing, you know, trying to find the needle in the haystack of all the crap that people are trying to sell you? You know, I would say when I started the company, it was mostly things we recorded, right? Because people didn't even realize that this was this was a product that they could sell. We didn't want to do oil yet. Right. And so, you know, you'd have a social media monitoring company who's used to selling Coke mentions to Coke and Pepsi mentions to Pepsi. And they say, well, what company do you want to track?

12:17And we said, all of them. And they say, wait, what? And we're like, all of them. We're like, that's not one of the options. We're like, well, then that's not helpful. So as people started to realize, oh, wait, there's value in a holistic view of all the companies that we're looking at at once, products started to emerge. Big financial companies have since then created a lot of products over time. So I think you're right. In the early days, it was really about recording and just having that asset. Today, there is a lot of incoming stuff. People are trying to sell you on different products. And I think needle in a haystack is a good way to think of it.

12:52I mean, some of it is going to be valuable and innovative. And you have to use your priors. You know, what's novel here? How is this different? and what's the edge to try to decide what you want to spend your research time studying, right? I want to study everything. I just have limited resources, so I have to pick and choose. And that's what makes a good researcher, somebody who has a nose for where to look for the next great idea. You referenced earlier the idea that the closer you get to the raw data, the raw capture, the more valuable that data is. But I also imagine that there is a hierarchy of features and like derivative features and that kind of thing.

13:31Like, how do you think about, how do you think about all that? Like, are you trying to do that? Well, first of all, is like, is your team focused on derivative features or do you just hand data over to modeling teams and they generate whatever derivative features they want? And like, how does that all come together there? Our team actually is also a modeling team. So we're both creating the features, but then also making predictions off them end to end. There are other teams that focus a little bit more on more differentiated techniques in machine learning that tend to take the features we create as inputs.

14:05But we do tend to do end-to-end work. In terms of the hierarchy, it's a great point. So when I said raw data is useful, I meant in the sense that if you have raw data, you can create any of those derived features in the future. Whereas if you deleted your raw data and you made derived features, and then you had a new idea, sorry, your idea, right? A great example would be you could record, you could look on television stations and you could record the words being said and then throw out the videos. And then one day you have technology to look at videos, but it's too late, right? Whereas if you recorded the actual videos, that's waiting for you as a future creation and avenue.

14:47On your hierarchy point, yes, absolutely. raw data is useful because it can create more things, but it's not useful in the sense that if you have 25 people that want to use your data, giving them a pile of raw bytes is probably not the most effective way to get them to use it effectively, right? I talked about, I don't know, looking at television stations. Let's look at an example of, I don't know, a CEO being interviewed on CNBC as an example. So I could just say, all right, everybody here has access to a video that we have and be done and everyone, or I could say, yeah, not just a feature, right?

15:26It's the ultimate mega feature. Yeah, it's not a good feature. And so then the question is, what are the features that you would want to create and share and pass on from that? And that's where the job gets really exciting. I mean, you asked sort of what the day in the life of our job is. I think the most exciting thing is to convert that data into hypothesis-driven ideas. Examples might be - The number of times they touch their nose. Exactly, right? Are you literally capturing that? You know, well, if I'm not, I can go back and add it, right? Because I have the raw data. So that's the idea. It's like any idea you have, you could go back.

16:02We could talk about how we would do that in a bit. But any interesting, how many times they're blinking, are they smiling, right? Right. There's obviously the words being said, which have a lot of value. But if you want to get even more wild, you could look at other things. How many questions were they asked? And you could even take a baseline of a CEO. The last 10 times they were on, they were X. What is, you know, the number of times the moving average of their nose touches, touch their nose is five, like readily. it. And, you know, I'm only half joking, right? Because it is possible that nose touching is nerves and that when somebody is more nervous, there's something else happening.

16:40And thus, you could. And so if I believed in that, if I did, I could spend my time, I can make that feature and I could go study, assuming I had history, a lot of assumptions. And I could say, all right, let's take 20 years of these interviews. Let's count nose touching and let's go study it. And any one of those things could be like a cool PhD thesis. But, you know, there's a lot to do. And So we're not going to that depth at all times, but those are the kind of ideas that we love to chase. Ultimately want to dig into like how the way you approach the, you know, capturing these features and analyzing and modeling off of these features has evolved with Gen.AI.

17:16But maybe before we do that is a good time to talk about how you might go back and pull out these new features. Maybe I want to say, is the person talking about the future or the past in these interviews? Interesting feature, right? What percent of their conversation is forward-looking versus backwards-looking? Cool feature, right? Historically, I might make a dictionary of forward-looking words. I can find past tense words. I could count. I can make a ratio. You don't need machine learning to count previous words versus forward-looking words, right? You make a ratio. It's kind of simple, right?

17:52Right. And that's sort of, you know, the one hot encoding NLP world of a decade ago. And even then it was cool, right? Like, like that, that is a real, I actually believe that could have value. And so even though it's simple, it doesn't mean it's not, it's not helpful. Now let's get to the, let's get to the nose touching. All right. Could I have done that 10 years ago? I think I could have, it would have required a lot of investment, right? It would probably be a six month vision problem, or maybe, maybe I'm, maybe it's three months. I don't know. And that technology was there 10 years ago. And if I really was like, you know, nose touching, this is my, like, I'm so confident that that's the best thing, that I could take all my resources and I could put it on nose touching and I can make a feature and study it.

18:36The problem with nose touching is that like, it's a cool idea, but is it worth six months of a person's time? Maybe not, right? And what's so cool about this new era of LLMs is that that six months might become six minutes, right? Hey, you know, did they touch their nose? how many times, you know, the profound shift in the ability to study the ideas in your head and the trade-off, the return on your investment, the ROI of these ideas, it's like that investment has been cut by 99 % for some tasks. And so all these things that were on your list, the nose touching stuff on your list that you're, yeah, one day that sounds cool.

19:15You know, you have this notebook of stuff that's just too hard. And all of a sudden you're, well, wait, wait, this isn't too hard. wait, this isn't too hard anymore. This is great. And so it's an exciting time because it used to be you felt prohibited by the tech, you would chase the things that were technologically within reach because that's a good use of your resources, right? And unless you had very high priors, you couldn't do more, but it was a big decision. And because it's no longer a big decision, it's awesome because it's like, it's kind of like a kid in a candy shop. You're like, I've had all this stuff and suddenly it's all in front of me.

19:49And LLMs has profoundly created this world where you don't need a team anymore. You don't need a giant project to answer basic questions like that, which makes the ability to create these features from either your head, you know, from the moment you think of them to the time you can analyze them, it's just a different world. And so in some ways it's kind of like a renaissance of feature creation, right? Out of the back of these LLMs. Can we go through an example in a little bit more detail and talk about some of the kind of technical bits of how you approach these problems? Look, like today's LLMs are multimodal.

20:26So you can give them videos and ask questions, right? So it's almost sadly, it's almost unexcitingly un-technical, right? What can you do? I can do a for loop across a thousand videos and ask, ask how many times did the person touch their nose? And I could put it to a CSV file. I know this is the simplest thing. And then, then, then I have the data and now there we go. So that's almost what's so wild about it. It's like, you don't, you don't even need to know anything about LLMs to imagine the ability to do this at scale in a new way. And that's kind of, that's what's really changing. Like, you know, five years ago, I'd have to tell you a diagram for you about how you do this all.

21:05But, but it's, it, that's, that's what's so profound is almost that it's actually kind of right in front of us. Yeah, clearly a big change is just the introduction of LLMs and your ability to query these historical datasets and create entirely new features, you know, from existing datasets. What else, you know, is evolving in the way you approach things? It's funny because in some ways everything's changing and in some ways nothing is changing. And what do I mean by that? As I pointed out earlier, we might've had earlier past tense, present tense ratios, right? We had one of my favorite things I've had, you know, there was a lot of dictionary work.

21:50I've got a dictionary of uncertainty words. I've got a dictionary of positive words and negative words. And I remember looking at a plot where uncertainty just spiked in the month of May. It's off. And I was like, oh, the word may. Oh, may. Yeah, yeah. Oh, you know, that sounds like the one hot encoding world can end up, you know, end up in a funny place. So I think the evolution that's really neat is in that world, you have these sort of the word, the word, you know, innovative and creative are just two words that have nothing to do with each other in the old world when I was in school. right?

22:29They're just two words floating in space. It's called one hot encoding because it's a big vector and every word exists as a one and a bunch of zeros, right? And you don't know anything about it. And what we've seen is this, this move to embeddings, which is really neat where instead of words just kind of existing in isolation of each other, they exist kind of floating in space, like near words that are like them. So cat and dog are kind of similar and, you know, creativity and innovation are kind of similar. And what that lets you do is really generalize your model because now what I learned about creativity, I learned a little bit about innovation, right?

23:06Those words kind of coexist in this cloud. So if I learned something about this area of the space, right, if I know what's happening in one quadrant, then I kind of know what's happening in all the words in that quadrant, right? It's hard to think in 300 or 600 dimensions. So I tend to think in three, you know, I'm feeble that way. But that really was sort of the underpinning of the movement where things started to get more powerful, where you weren't just counting words anymore. And I think what people don't realize is that, going back to my syntax versus empirical stuff, is that LLMs actually, they're starting to add reasoning to them.

23:47But certainly as of a year ago, the chat GPT we were all introduced to, it was just trying to imitate humans, right? It wasn't knowledgeable. It didn't say, you know, cats, the dog barks because it knows the dog barks. It said the dog barks because, you know, a logical thing after the word dog is the word barks. And it was wild to feel like these things know about the world when, in fact, they're just kind of parroting humans in this really interesting way. You know, this idea of taking advantage of embeddings and kind of the semantic relationship between words as being a big shift in the way you're able to approach problems.

24:28Are you getting that kind of for free passively because it's part of this whole LLM world? Like it's part of the way LLMs are built and operate? Or are you like directly, you know, building things around embeddings? Like I remember for a while, you know, clearly it's still, it's increasingly popular for organizations to build their own embeddings and to like create embedding spaces for RAG types of applications, that kind of thing. But I remember conversations, you know, like even five, six years ago, maybe I'm thinking of like a interview with Instacart. Like they built an embedding around like groceries and stuff like that.

25:09Yeah, right. Exactly. Everything's a vector. Everything's a vector. Exactly. Just walking vectors. And so are there use cases in which you're doing that directly as opposed to, you know, just... Absolutely. Talk a little bit about those. I mean, ultimately, our goal is to make predictions about where, you know, say stocks are going. So if I'm looking at news stories about a company and I want to make predictions about what that means for that company, right? I guess I have to build a bridge between that news story and some algorithm that's going to make a prediction. And whereas in this one hot world that was, there were methods, it is a much more rich world to be able to convert that story into an embedding or multiple embeddings or, you know, there's different ways you could approach that.

25:59Once you're in this vector space, it becomes a much more natural connection into the machine learning world, right? If you've got, if everything's a 300 dimensional vector instead of a blob of words, you know, What comes downstream is, and people can imagine, is various sophisticated learning algorithms that can try to tie the thing that you're embedding in whatever creative way you want to into a prediction about the world. So it's kind of like you were saying before, like you alluded to like same, same, but different. It's like you are doing many of the same things like sentiment you referenced, maybe, you know, related companies, related people, building knowledge graphs, all that kind of stuff.

26:43And now you just have a more powerful way of doing it that is semantically rich as opposed to based on word counts and that kind of thing. It's more, it's more semantically rich. And so those vectors themselves could be features, right? They're, they're, they're, you can make features about a company. You can make features about a story you can make. So, um, and the LLM, I mean, under the hood, it's, it's embedding things. And so if you have, if, if you have access to the internal workings, you can also pull, pull the, the embedding right out of the, right out of the process, depending on, on what the software is.

Read the full transcript

27:17Um, or you could take the output and, you know, do what you, embed that, right? You, Like you can do so many different, there's so many ways to interact. So I think the very simple way is, hey, LLM, give me the sentiment of this thing. And it gives one through 10. And like that, I could do that. That's very straightforward now. That wouldn't have been straightforward, obviously, years ago. But that's the thing about financial modeling is that we, it is so incredibly noisy, right? In machine learning, I think people don't fully think about the fact that in most machine learning tasks, there is an answer that you could get given the information you have.

27:55So if I had to say, hey, is this a cat or a dog? You know, you could figure out the answer. Usually, hey, you know, what should I do in this car? Should I slow down or speed up? There's usually a good answer. There might be like something crazy happens or not always. But in finance, you actually, it's not clear that no matter how much information you have, that there is enough information to predict with any real certainty what's going to happen tomorrow, right? And so we're making this prediction to the future, which means that the signal to noise ratio, the noise is just so much bigger, right?

28:31And so if we can predict an R square, if you see an R squared of like 0.05, you know, something's broken because that's way too high. Like, you know, like there's no, these things that you're trying to predict are very small signals, right? the nose touching might be real, but it's not gonna be predicting most of that stock price. It's gonna be a tiny, tiny, tiny, tiny, tiny portion. And so, yeah, that's one of the challenges of the financial markets. And also - And so from that perspective, the signals that you're operating with are so noisy. The fact that the LLM, the fact that the difference between a five and a six and a seven sentiment is kind of random, doesn't really matter.

29:15Maybe, right, yeah. Or it matters very, very little in a way that's hard to leverage. You know, it's hard to tell. And so are these, we talked a little bit about the hierarchical nature of these signals. I'm curious about how you put that all together. Like I'm envisioning, you know, at a certain point, And, you know, particularly if you've been doing this for 17 years, like you've collected a crap ton of data, like you're trying to update, you know, a large portfolio of features across that raw data. You're trying to enforce some degree of consistency across that. Like I start to think of like, you know, what platform, you know, what the platform looks like and what the tooling looks like and that kind of stuff to do that at scale.

30:08It's funny, so one thing about Two Sigma is that we are a, we call ourselves a platform company because ultimately, you know, it's a very tech heavy investment company. And we have a lot of different teams looking at a lot of different types of data, but they are doing that together into sort of a centralized set of portfolios. some other financial companies have these competing pods that sort of fight each other, you know, to try to have the best performance. And we're not like that. We're all centralized. So we have a shared, we have a shared platform that, that kind of, we all can feed on. And so what does that look like?

30:50We try to break up the investment process into abstractions that have, you know, good APIs, right? So I talked about this model is something that predicts something, but in terms of a platform, right? What does that mean? It means I have a set of inputs, right? Which are going to be timestamped information. And the API, well, my job is to, you know, send out a prediction for every time cycle, what I think is going to happen to everything that I'm looking at, right? And I'm going to do that in a loop. And so if someone, if you provide a modeler with access to data and this API on the output of what they can predict, then their job is to connect the raw data or the features that they've read in to that prediction.

31:32And what happens in between those two points is sort of the magic of that particular model or the hypothesis or the machine learning or whatever's happening. Obviously, you need a lot of storage. You need good search. And a lot of tools to run these back tests are very complicated. If you're trying to predict what would happen had I made this decision over the last 15 years on the market, there's a lot of technology that goes into replaying the set of stocks and what would have, you know, creating a virtual world where you're backtesting this new idea and see how it would have done. So that's all kind of built into the platform where someone likes me, I can just call, okay, here's my forecast.

32:13Go tell me how I did. And I can put that on the, I can put that on the platform and it can come back with the measurements for me without me having to think about how or what's happening. So those abstractions are really helpful as, because, you know, the process of turning raw data into a portfolio has many, many steps. And, you know, my team's responsibility is turning raw data to features, then features to just predictions. And after the predictions, another team is thinking about how to make all those predictions, combine them, and build a portfolio for our company that's going to maximize the returns for our investors.

32:49And so the distinction there is that a prediction that your team is making is whether, well, are you predicting the future of the future value of features or are you predicting the future of an equity price based on features? It sounds like there's a line somewhere between one of those things and predicting, you know, the value of the portfolio given a decision to buy or not. So we're predicting the future price of, if we're doing equities, which is where I tend to focus where I focus and where I work. We are predicting the future price of equities that we trade. Based on a feature set. Based on everything you have.

33:38Based on everything we have. Yeah. And so our role is to build a bunch of different models. One model might look at CEO interviews. One might look at fundamental prices. One might look at what an analyst is doing. And so you end up with literally thousands of these models all interacting and they all have their own view of the world. And the idea is, you know, to keep your signal working, they're hopefully orthogonal enough to one another that when one's not working, you know, and so that's how the company works overall. But the thing is, once you have all these predictions, you can say, all right, we overall, we think that this stock price is going up and this stock price going up.

34:14But how do you turn that into a portfolio that doesn't have certain risks that you're not to expose to a particular industry or you're not, you know, so you have to keep everything neutral so that if something happens in the market, we can stay balanced. And so, you know, that's not the part that I, that's called optimization. It's basically converting the predictions into a portfolio that's been optimized to be kind of balanced well. And so that's like its own research area, their own set of researchers, their own set of technology that then decides what optimal portfolio we want to be in at that particular time.

34:50I think chasing down a, the signal that I'm latching onto and chasing down here is you kind of anthropomorphize the idea, are these models having their own point of view or something along those lines, as you said. It made me think of some kind of agentic extension of that where the models are arguing with one another about the direction that they think something is going. Is that a thing that you're thinking about? I mean, if I wasn't, it would be weird, right? I mean, you're hitting the nail on the head. I mean, we have these different views of the world that have been thoughtfully created by people with hypotheses, with views, encapsulating from a set of features.

35:37You know, some might be very interpretable and simple. Some might be more complex in a black box machine learning way. You know, we always have humans on the end to monitor everything. But ultimately, we have all these different models with all these different approaches and techniques. And yes, do I think a lot about what an agentic approach to that looks like? Like, absolutely, I do. It's terrifying and exciting all at once. But what's your best guess today at what an agentic approach to all that looks like or means even? My best guess is that you have different agents that have, I mean, this is sort of the point of these agents, right?

36:12They have different understanding. They're experts on different things, right? One is in depth, understands currencies and one understands what the Fed is saying. And, you know, I think ultimately their ability to communicate with each other and say, all right, well, I know that this company has, you know, I know that interest rates are important for this company. Let me go ask the interest rate expert what's going to happen. And then I'm going to go ask the, you know, the weather person what's happening tomorrow. So I, and that's, this world of these, these, these, these agents making guesses about the things that they're experts on.

36:44And yeah, I don't exactly know what it's going to look like, but definitely an area where we're experimenting in because it clearly, look, any of our roles feel like they're heading towards automation, right? Many, many, many, and I don't think just because it's research or innovative, it doesn't mean that a good reasoning LLM can't start to fill in some of those gaps. And I find it, yeah, I find it pretty fascinating, but also exciting because, hey, can I have 10 ,000 researchers at my fingertips? That would be pretty powerful. And in some ways, it feels like that's the way we're going and faster than I would have guessed.

37:26One thing I'll say is that what I've learned is that innovation cycles seem to be happening pretty quickly at an increased pace. You know, it is business as usual. NLP has always gotten more advanced over the years, but it feels like the cycle is speeding up. And so you need to be ready with a platform that can adapt nimbly to the latest technology, right? So if you're going to go spend a year training up your own thing and then you come out for air and the technology you use that you started a year ago is now obsoleted, you know, that's not going to work. So you really need to be front footed and look at what's happening in the world and build on top of the ability to take in modern open source approaches or work with vendors and tools that are adapting very quickly.

38:14And your platform needs to be able to jump on and off different things at different places to stay up with this new innovation cycle because there's no time to stop and build something in this space that's going to be able to keep up with what's happening. And so that's kind of fascinating as well. So interesting. Yeah, you mentioned open source and you kind of started to allude to like how you think about build versus buy. I'm sure it's an oversimplification to say that things are moving too fast for you to build anything. Right, that's a oversimplification. I think the things that you build need to be plug and playable with the changing world.

38:54Can you give me a concrete example of navigating that? While there's a lot of interesting things you can do with, say, fine tuning in LLM or something like that, you want to be in a position so that you can swap in the next model, wherever it comes from, however it comes from, very quickly. So that the rest of your stack is not like, oh wait, it's a new model. It's from a new company. Everything now has to be redone, right? So you have to design your inputs and outputs so that you can wrap a new model around something such that it's easily ingestible into your system as opposed to picking something and then customizing on top of that one thing.

39:39Yeah, I'm hearing kind of like the minimum to build is like owning the abstraction. Right, right. So that you can swap things out and then maybe you're selectively, I mean, I've got to imagine that you're, you know, well down the road of fine tuning your own LMs for in areas that it makes sense. Yes? Yeah, yeah. Which, yeah, obviously that's an important aspect of it. But even with fine tuning, the things that you had to fine tune six months ago, you know, the next model somehow is already, you know, risks being better at the thing you fine tune your six month old model, right? Like, I mean, it's moving that quickly.

40:19And so it's kind of wild. You'd think that the level of depth that a generic LLM has on, say, financial accounting statements, I mean, it's incredible. It's in any field, right? We've all seen that. But so maybe I had to fine tune last year's model, but suddenly this year's model seems to have all that ability. So that doesn't mean that you're not fine tuning. It's just that the assumptions that you have about what you need to do might be invalid six months down the line as the technology starts to speed up, is my point. With that in mind, how do you think about where you make those build investments?

41:02like, and I'm specifically trying to get at, you know, it would be easy to say, like to try to project, well, you know, the LLM is going to be able to do this for us. And in fact, it's not just easy, like it's necessary in a lot of cases to say, yeah, I'm not going to, I'm going to force myself not to go down that rabbit hole because, you know, by the time we're done with X, the LLM is going to be able to do that, you know, But sometimes you sit around waiting and it doesn't actually fill that gap for you. How do you think about riding that edge? I mean, it has to do with your priors on how important a particular build is, like how fast you need to get it out.

41:42And if you think that you're in a position where you have an edge to utilize that novel thing you're doing, and especially in an industry like finance and quantitative finance where you're literally competing for prices. Like, you know, it's someone else's machine learning is against your machine learning and it's a zero-sum game, right? So, you know, time matters. If you're six months out the gate before the other group, that means that that can make a big difference because what you find, by the way, is when you trace signals over time, they get weaker and weaker and weaker as your competition starts to come in.

42:19And so it's called alpha. It's kind of efficient market. Yeah, it's an efficient market, right? Any innovation you think you have, well, now everyone has access to data X or Y. And so it slowly decays away over time. So, you know, there is a time component and you have to, it comes back to that return on investment question and your priors as a researcher of how important that particular thing is to get now versus later. You know, when I think about this, again, this kind of hierarchical, you've got these base level features, you're building higher level features. It made me think a bit of kind of the old ideas, quote unquote, old ideas around like AutoML, like a data robot thing, like I'm going to just take all these features and run some statistical programs across them to like combine them in different ways.

43:04Like, are you doing a lot of that kind of stuff? I mean, it's hard not to, right? You're staring at this pile of data. I think it would be, you know, I wouldn't be doing my job if I haven't explored what happens when I use them all. So, yeah, it's a complicated question because the risks of key hacking and overfitting, you know, when you start to take everything at the wall, throw it and see what sticks, it opens up new scientific rigor questions. And we live or die off of, you know, whether we overfit or not, right? If you make that mistake and your model doesn't trade, that's not going to go well for a modeler.

43:40You can have, you know. So we have to be very careful. But if you do it in a rigorous way and careful way, then yeah, I think there's a lot of exciting work to be done in automation end to end. And it's definitely a focus area that we think a lot about and is growing as technology grows. Yeah. In a lot of ways, it touches on a question that comes up in lots of different spaces around like modular versus end to end. So, for example, I've had this conversation with a number of guests in, you know, robotics, autonomous vehicles where like they're, you know, the industry has priors of like you've got, you know, a sensing system, a control system.

44:27you know, if you've got these subsystems and you want to wire them together in ways that make sense. But then there are, you know, that's kind of one pole. And the other pole is, well, yeah, let's just take off the data, feed it into a model. And, you know, if we have enough data at some point, you know, the model should be able to do it. Like, is that a tension that you have to? I mean, yeah, absolutely. It reminds me, again, back to my days in the PhD program where there was this sort of syntax versus enough data view. And there used to be this, competition for who can make the best, you know, machine translation system.

45:03And there were all these fancy algorithms. And one day Google entered and they had, you know, the five-year-old algorithm, but they had this massive model, like 10 times bigger, a hundred times bigger than anyone ever had. And they beat everybody. And everyone was like, wait, like we have all these fancy ideas. Yeah, but we have a lot of data. And it kind of changed, it changed the world. And so I think with enough data, you can start to do really interesting things. And so, yes, even, even I described to you these abstractions, I'm in modeling and my job is to turn data into forecast and someone else does that, you know, what is the word look like when you just hook raw data up to, you know, a portfolio and you, you, you start to take away those abstractions.

45:45What, what new things? And there's definitely a tension there. I think that if you, if you, if you force yourself into a set of chained events, you might lose a lot of the power of machine learning because at each handoff point, you've stripped things down into your assumption of what matters. And that algorithm might disagree, right? I mean, in fact, machine translation is a great example of that. It used to be voice to text, text to text, text to voice, right? So if I wanted to go from English to Spanish, it would go English, you know, I would say it, it would turn to text, then they would use a backend model, then they would generate.

46:21Now it's just, you know, they don't, that's not how these things work anymore. They've hooked the language. They know how to convert it from audio to audio. It's not running any text in the background. It's an exact example where if somebody had forced that hop, you would actually have a less good system than we have today when they said, oh, look, let me remove these abstractions that humans have added and just let the system go with enough data. And it turns out that can work incredibly well. Same time, I'm thinking of a conversation I had with Scott Stevenson, who founded a company called DeepGram that's in like the speech-to-text, text-to-speech base.

46:55And one of the things that he observed is that, you know, while skipping, while staying in the audio domain can be really powerful, sometimes jumping into the text domain is useful because as us humans, we can't look at the spectrogram and like intuit anything about it. But like text is useful for debugging. for understanding the system. Are there analogies in your space along those lines? Yeah, I think interpretability is an interesting thing. And as you go more end to end, you lose more of that visibility into the world, right? And I get asked a lot about how comfortable are we with black boxes in investing, right?

47:40These are big decisions we're making. And what I like to say to people is, well, have you ever gotten under general anesthetic? So yeah, I mean, when I got my wisdom teeth out or something, I said, okay, did you know that, you know, modern science doesn't actually know how that works? We've just tested the inputs, we've tested the outputs, and we've grown comfortable with enough data that general anesthetic has particular properties. And we, in medicine, they're comfortable enough to utilize it by just having very careful measurement on the input and output. And so I see, I love interpretability.

48:11I think I'd rather, by the way, have a medicine that people understand. Don't get me wrong. I think that's a more comforting thing and you'd have higher priors. But if you study something enough, that interpretability maybe isn't as important and maybe our ability to understand, it might not even, even if it could be explained, it might be that the thought is so complex that it would be hard to explain even in the most perfect world. And so I tend to prefer interpretability. I have higher priors when I understand what's happening and it's all about being equal, like Occam's razor, simple is better.

48:46But that doesn't mean that we should be so naive that we can turn away anything that we don't understand. At the risk of bouncing all over the place here, I remember you mentioned open source and kind of going back to that platform conversation. I'm curious, can you articulate like what, you know, like, are you using one of everything or are there specific, do you have strong feelings around like, you know, O3 is, you know, the state of the art for us and we're doing everything with XYZ or we've hard pivoted to Llama 3 and use it for everything like, or, you know. Yeah. So, you know, one of the interesting things about these models is, remember how I told you that we're terrified of temporal leakage, you know, timestamps matter.

49:38so it is kind of scary for us to use a a you know off-the-shelf model that's been trained on data in 2020 to ask questions from a document of 2019 right so if i could say hey here's this enron conference call you think it's good or bad is the word enron going to trigger a negative reaction because you know somewhere deep in the psyche of lm and knows there was a big bankruptcy and so you have this sort of sneaky look-ahead bias that could kind of type in. So, you know, in the build versus buy world, there are risks to using these off-the-shelf things because you can't control the data that's going in and you might fool yourself.

50:18And so that's one reason why we do prefer open source more controlled experiments because we know what went in and we know that what we're seeing we can actually believe in. In terms of the best model, by the time I answer the question, it would probably be different, right? So I guess my quick answer is no, we're not, it's not the kind of thing where, I mean, look, I have opinions, this is working out better than that at this task and that task. But nothing that, it's clear that this is going to be a iterative process where different, there's going to be different advantages to different models are going to be hopping over each other, it's going to be happening rapidly.

50:56And there's not going to be a horse to bet on, you're going to be well suited to have a diversified set of inputs to, you know, build interesting and orthogonal outputs. And you need to be comfortable using a wide array of technologies, not just kind of betting on a single one. But can you say something to the effect of like, given your preference towards open source and understanding the, you know, training environment, data recipe, whatever of a model, you know, you've been able to shift, you know, you know, some significant percentage of your workloads to open source models or, you know, conversely, you know, the model yet another feature in this basket of features and you train all of your data on all of the models and it's just, you know, one level higher that you go to make sense of things.

51:46Yeah. I mean, I kind of, if I could stop my fingers and train all my data on all my models, I think I might. I mean, it's nice to have it all there, right? Then I wouldn't have to run any experiments. So I love that concept. And I think Like, actually, a colleague of mine last week was framing it basically as combinatorial research, right? Yeah, exactly. Right. The ultimate scary place to go and not get buried. And so, you know, I think there's real truth to that. I do think that when you're making a prediction about what's going to happen, much like we have different models that combine to make Two Sigma's forecast about what's happening next, right?

52:24the model looking at this data, that data, so too is their value in having an ensemble of LLMs that each have their own strengths and weaknesses. And it's not, I'm not always, I'm not looking at the best at things. I'm looking for, you know, a group of things that each have their own take that when I average out among them, I'm better off and more robust in the future than had I just picked one. So yeah, I'm not, I really am of the belief that there are pros and cons And that in itself is a strength because it builds orthogonality, which when you combine orthogonal signals, you get a much smoother response than when they're correlated signals.

53:05And you kind of said something that alluded to, you know, just like, you know, all your data is timestamped, you know, maybe all of your models are timestamped and you have like you know, a model with a 2019 perspective or a model with a 2020 perspective, you know, etc. Are you, you know, and again, speaking of combinatorics, like that explodes really quickly. It's yet another dimension. It's yet just, yes, another dimension to add to the factorial. Yeah, it's very, it's, you're exactly right. And so, but it sounds like you are doing some of that. Yeah. I mean, ultimately, look, we've always done that, right?

53:47So if I'm training even a linear model, there's different ways you could do it. If I'm trying to predict the nose touching as a predictor of a company's performance, right? I could, if I take my entire sample and I calculate some beta between the number of times they touch in the company performance, then I guarantee, I just overfit my data and it will work on my, it won't work in the future, but it'll certainly work on the history, right? So what's a better way to do it? The better way to do it is to kind of think about this temporally and say, okay, well, if I had the data from 2000, 2010, what would I have done in 2011?

54:29And, and then, you know, so we, those, those are some small changes you can make to, to avoid. Now you're still fitting because you still have to be careful. It doesn't solve all your problems, but, but yeah, we've always lived in a world where, where you have to be very careful to not get future data in, in your experiment. And the thing about complex machine learning is if you happen to give it a leak into the future, it will find it, you know, like a linear regression might not find it, but a, a neural network that sees, oh, wait a minute, I already know this. you know, it's going to pull whatever weird hole you might have accidentally given it and pull hard and, you know, and be successful.

55:05So I think time is so important. So what measures do you take to prevent that kind of data leakage? We're pretty strict about timestamping our data at the company. And so I say all data is a time series. People say, what are you talking about? I'm like, well, every piece of data has a timestamp on it. Well, what do you mean? You know, I mean, yeah, tell me, well, you know, and if you think about it, all data came to exist at some point. And so you can inherently add a timestamp to almost any data. And if you do that, and if you're careful to make that data, that timestamp, the creation date of that data, then you can easily query, hey, I want to join my information, but only backwards in time.

55:48And so if you constantly are only looking backwards in time, then you kind of save yourself from that leakage. Now there are risks, because vendors don't always, you know, abide by the same rules. For example, let's say a company wants to track the S &P 500 stocks and they're going to give me, you know, the website job postings of S &P 500 and they come and they go, oh, I have all this data on the S &P 500. So they give me a 10-year dump. But the problem is if they give me a 10-year dump of the S &P 500 today, and then they give me 10 years of history, that has look-ahead bias because 10 years ago, we didn't know who the S &P 500 would be in 10 years.

56:26All I have to do is buy all those companies for all 10 years and I will make a lot of money. But I think that I want to do every time someone hire someone just buy stock and I would make so much money. Right. But it wouldn't be real. I wouldn't generalize to the future. Now, the vendor is not doing this because they're trying to trick you. They just don't think temporally in that same way. And so, you know, you do have to be very vigilant in the timestamps people put in and to make sure that that leakage doesn't come because again, in the world of modern machine learning, that leakage can be, if you think that you have this model that's so performant and you're being tricked, then, you know, it's not going to, that's not going to show up in, in its live trading.

57:03And then you're, you're going to notice that and eventually turn that off, but it's going to be a waste of your time. So thinking about forecasting and predicting the future, what do you predict for the space? And what are you excited about, you know, coming down the road? I got to say, I think your, your question on the agentic stuff is quite exciting. It's hard to, for all of us in many fields, to wrap our head around what's gonna happen in the next six months to a year to two years, or five years down the road.

57:35I'm excited about the increase in efficiency that brings and the ability to convert ideas into signals at a faster and faster pace because, you know, then you can just take your most creative people and create data out of them, you know, at such rapid speed. And that's really exciting. It opens up so many different avenues. And of course, it opens up challenges, right? As these tools, as the LLMs get stronger, if they're able to code as good as, you know, your junior engineers or able to do research as good as your, you know, what does that mean for the future of a company like ours and who we hire and training and things like that?

58:16So I think that's the challenge, but I think that the upside outweighs some of those challenges in that the opening up of data is just so profound and exciting that that keeps me excited to come to work every day. Well, Ben, thanks so much for taking the time to jump on and share a bit about what you're up to and what you've learned. Yeah, I enjoyed it. Thanks for the chat today. Yeah, thank you.

58:48Thank you.

From the publisher

Today, we're joined by Ben Wellington, deputy head of feature forecasting at Two Sigma. We dig into the team’s end-to-end approach to leveraging AI in equities feature forecasting, covering how they identify and create features, collect and quantify historical data, and build predictive models to forecast market behavior and asset prices for trading and investment. We explore the firm's platform-centric approach to managing an extensive portfolio of features and models, the impact of multimodal LLMs on accelerating the process of extracting novel features, the importance of strict data timestamping to prevent temporal leakage, and the way they consider build vs. buy decisions in a rapidly evolving landscape. Lastly, Ben also shares insights on leveraging open-source models and the future of agentic AI in quantitative finance.

The complete show notes for this episode can be found at https://twimlai.com/go/736.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
LLMs for Equities Feature Forecasting at Two Sigma with Ben Wellington - #736The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 1 h
Listen in VO