In short
Whether large language models (LLMs) can improve time series forecasting, and when they should be used versus traditional numeric models.
Guests
No guest names or backgrounds are provided in the transcript.
Key claims
LLMs outperform traditional unimodal time-series models when data is “shifting” (distribution changes unpredictably, e.g., COVID spread, exchange rates, erratic weather) or “high transition” (frequent complex state changes, e.g., high-frequency trading, power grids in storms). A major requirement is pre-alignment: translate numeric time series into the LLM’s language-compatible representation; post-alignment (joint fine-tuning) causes catastrophic forgetting/pseudo-alignment and fails in most tasks.
Notable examples
8 billion observations across 17 scenarios and four horizons; routing analysis shows the system sends tokens through the LLM only when shifting/transition is high, reducing errors.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOSkepticism vs Optimism in LLMs
1:07 to 2:10
Examining the divide between skeptics and optimists regarding LLM capabilities.
“Yeah, a lot of people think it's just hype.”
Traditional Forecasting Models Explained
2:10 to 3:31
Understanding traditional time series models and their limitations.
“And it covers everything from in-domain data to entirely unseen out-of-domain data.”
LLMs as Well-Read Historians
3:31 to 4:51
Contrasting LLMs' contextual knowledge with traditional models' lack of context.
“It just sees a sequence of fluctuating values.”
Cross-Domain Generalization in LLMs
4:51 to 5:46
Explaining how LLMs can apply knowledge across different domains effectively.
“Okay, I understand the historian has broader context, but a localized server outage isn't a paragraph of text.”
Pre-Alignment vs Post-Alignment Methods
5:46 to 11:03
Discussing the methods for aligning LLMs with time series data and their effectiveness.
“it learned in one domain, like language or economics, and successfully apply it to a completely different unseen field, like server traffic.”
The Importance of Pre-Alignment
11:03 to 11:41
Highlighting how pre-alignment is crucial for successful LLM predictions.
“You induce what's called catastrophic forgetting.”
LLMs in Practical Applications
11:41 to 14:00
Analyzing when to use LLMs based on data characteristics and scenarios.
“But wait, if we assume we have the perfect pre-alignment translator built, let's say that's totally solved, do we actually need to run every single piece of data through a massive computationally expensive LLM?”
The Role of LLMs in Chaotic Data Forecasting
14:00 to 18:01
Explore how LLMs enhance forecasting in rapidly changing environments.
“Imagine multi-stage dynamics where a system is constantly evolving through different phases very quickly.”
Understanding Model Limitations and Requirements
18:01 to 19:26
Learn about the limitations of LLMs and the necessary conditions for effective use.
“Scaling up the model size does not automatically yield better forecasting.”
Essential Questions for Implementing LLMs
19:26 to 20:35
Identify key questions to address before using LLMs in data analysis.
“is really the only way to successfully scale these models for forecasting.”
Show all 11 chapters
The Potential of LLMs Beyond Traditional Data
20:35 to 21:48
Consider how LLMs might apply to unexpected fields governed by patterns.
“So if you were sitting in a strategy meeting next week, and someone enthusiastically suggests, hey, let's just throw an LLM at our supply chain data to see what happens, you don't just nod along anymore.”
Transcript
Automatic transcript. May contain errors.0:00Imagine taking the entire works of Shakespeare, every Wikipedia article ever written, and I don't know, a decade's worth of Reddit threads. You feed all of that into a computer. Right. A massive amount of text. Yeah. Just an ocean of text. Yeah. And then you use that machine to perfectly predict tomorrow's stock market crash. I mean, it sounds completely absurd, right? Oh, it sounds entirely like science fiction. Exactly. Because language AI is supposed to, you know, draft our emails. Maybe you write some code or plan a vacation itinerary. It is definitely not supposed to run our spreadsheets.
0:34No, traditionally that is completely outside its wheelhouse. But whether or not large language models or LLMs can predict the future of cold, hard numbers is actually the absolute frontier of tech right now. And honestly, it's a frontier that has a lot of people highly, highly skeptical. It really is the defining technological battleground in data science today. I mean, we are looking at a landscape where the tech and business communities are just fiercely divided on what these models can actually do with numbers. And that divide is exactly what we are jumping into for this deep dive. Because on one side, you have skeptics pointing to recent studies, usually smaller ones, claiming that throwing a massive language AI at numerical data is just a colossal waste of computing power.
1:19Yeah, a lot of people think it's just hype. Right. But on the other side, optimists believe LLMs are the ultimate key to unlocking really complex real-world data predictions. We're talking about time series forecasting here. Which is essentially the science of predicting what happens next based on historical data. Right, so things like stock prices, traffic patterns, sudden weather shifts, or supply chain metrics. And to settle this debate today, we're looking at a massive new study. It's titled, Rethinking the Role of LLMs in Time Series Forecasting. And I have to say, the sheer scale of this source material is really what makes it definitive.
1:53The researchers didn't just look at a handful of spreadsheets to test this theory out. No, they went huge. They analyzed a staggering 8 billion observations. I mean, they tested models across 17 different forecasting scenarios and four distinct time horizons. That is just a massive amount of data. It is. And it covers everything from in-domain data to entirely unseen out-of-domain data. 8 billion observations is serious weight. But, I mean, I really have to push back on the fundamental premise here early on. Okay, let's hear it. Because if you're listening to this and wondering how this even makes sense, you're not alone.
2:29Why would an AI trained on human language be any good at reading numbers? It's a very valid question. And LLMs is actually a next word predictor, right? It's not a calculator. How does knowing the linguistic structure of a poem translate to predicting a sudden drop in a localized supply chain metric? It just doesn't connect for me. Well, that skepticism is exactly where the researchers started. And to answer it, it requires us to look at the underlying mechanics of how these models actually think. Let's look at traditional forecasting models first. Okay, lay it out for us. Traditionally, time series models are what we call unimodal.
3:06They are built from scratch and trained purely on raw numerical data. Right, so a traditional model only sees the math. It sees a line on a graph going up, and it runs a statistical formula to predict if it will keep going up based on the historical angle of that line. Yes, but those numerical representations are completely abstract. I mean, they lack any explicit encoding of the real world context. So they don't know what they're actually looking at. Exactly. A traditional model doesn't know what a stock market actually is or what a hurricane is. It just sees a sequence of fluctuating values.
3:39Okay, let's unpack this. I like to think of a traditional forecasting model as a brilliant but incredibly sheltered mathematician. Oh, I like that analogy. You lock them in a basement, you slide historical sales numbers under the door, and you ask them for tomorrow's forecast. They might find some hidden mathematical patterns, but they are entirely blind to the outside world. Right. They have no context. They don't know there's a global pandemic happening or that a major shipping channel just got blocked by a massive cargo ship. They just crunch the numbers. That is a highly accurate mental model.
4:12Now, contrast that sheltered mathematician with an LLM. And LLM acts like an incredibly well-read historian. Because it's read basically the whole internet. Exactly. Because it was pre-trained on massive text corpora. It possesses what the study calls pre-trained world knowledge. Pre-trained world knowledge. Yeah. It inherently understands the semantic relationships between, say, a shipping lane blockage, the subsequent supply chain delays, and the resulting economic impact. Because it's literally processed millions of texts detailing those exact human behaviors and economic systems. Precisely.
4:48It knows how the world works, not just how numbers trend. Okay, I understand the historian has broader context, but a localized server outage isn't a paragraph of text. It's a sequence of numbers, like a sudden spike in ping times. Right. So how does knowing the history of the 2008 financial crisis help an LLM predict server traffic? That still feels like a massive leap. Well, it comes down to a really fascinating concept called cross-domain generalization. Cross-domain generalization. Okay, what does that actually mean in practice? Underneath all the text, an LLM isn't just memorizing words. It's mapping the structural patterns of complex systems.
5:23Okay, so it's looking deeper than the vocabulary. Exactly. The way information cascades through a social network during a crisis has a remarkably similar mathematical shape to the way a bottleneck cascades through a server network. Wait, really? So the underlying structure is the same? even if the subject matter is totally different. Yes. Cross-domain generalization is the model's ability to take the underlying structural logic it learned in one domain, like language or economics, and successfully apply it to a completely different unseen field, like server traffic. So it's identifying the universal rhythm of chaos, regardless of whether that chaos is written in English or in server pings.
6:02That's a beautiful way to put it, and the data supports that entirely. The study demonstrates that models trained on multi-source time series data, really leveraging that deep LLM world knowledge, consistently beat single dataset traditional models. By how much though? Are we talking a minor edge here? Not at all. They beat them in over 70 % of in-domain tasks, and they also exhibit significantly stronger out-of-domain generalization. Meaning they can predict patterns in industries they've never even seen before. That is wild. Wild. Yeah, it's a huge leap in capability. But wait, if the numbers are that definitive, why were the skeptics so convinced LLMs were useless at this?
6:44I mean, where did the whole anti-LLM narrative come from if the data is this strong? Well, the skeptics were evaluating LLMs under highly limited settings. They tested them on really small, single-task data sets. Oh, so they weren't giving them enough to work with. Exactly. And mechanically, they were often only utilizing the shallow layers of the LLM. If we go back to your analogy, it's like hiring that brilliant, well-read historian and only asking them to balance your checkbook. Right. If you don't give them complex problems, they just look like an overpriced calculator. Yeah, you completely fail to utilize their true multitasking power.
7:18All right. So the historian beats the isolated mathematician. I buy the theory. But if the LLM is built to process language, how do we mechanically get a billion rows of an Excel sheet into its brain? That is the big question. Because you can't just copy-paste a massive database into a chat prompt and expect a nuanced forecast. It would just break. No, you absolutely can't. And you've hit on the modality gap. We are dealing with two completely different types of information architectures here. Right. Miracle time series data on one hand and semantic language data on the other. Bridging that gap is the single biggest technical hurdle in this whole field.
7:55It's known as the alignment problem. And the researchers rigorously tested the two main paradigms for solving this, pre-alignment and post-alignment. Okay, pre-alignment and post-alignment. Let me make sure we are translating this into something tangible for our listeners. Go for it. Let's say the LLM is the CEO of a multinational company, but the CEO only speaks English. Okay, I follow. And the data coming in from the field is a highly complex financial report written in, say, specialized accounting shorthand. Is pre-alignment essentially hiring a dedicated translator to format and translate that shorthand into clear English before ever handing the report to the CEO?
8:34That captures the dynamic perfectly. Mechanically, in pre-alignment, you map the time series data into language-compatible representations before it ever touches the LM. Before it even gets to the CEO's desk. Exactly. And the study details doing this via a process called cross-attention with word embeddings. Okay, that sounds like heavy tech jargon. To define that for our business listeners, cross attention with word embeddings basically means taking the numerical data and mathematically matching its shape to the shape of words the LLM already understands. That is a great way to put it. You are using a separate, smaller neural network, your translator, to project the raw numbers into the structural language space of the LLM.
9:13Right, so the numbers start looking like a language. Exactly. And crucially, you do this without altering the LLM itself. The CEO's brain remains untouched. Which brings us to post-alignment. If pre-alignment is hiring a translator, post-alignment must be forcing the CEO to sit in the boardroom with the raw accounting data and making them learn the shorthand on the fly. It's even more invasive than that, honestly. In post-alignment, you are jointly fine-tuning the time series encoders and the LLM simultaneously. Oh, wow. So you're messing with the core model. Yes. You are actively updating the foundational parameters of the LLM using supervision from the forecasting task itself.
9:50That sounds incredibly risky. I mean, to push the analogy, that's not just making the CEO learn shorthand. That's forcing the CEO to rewrite their entire foundational business philosophy just to understand one specific accounting report. That's a very real danger. It seems like it would make them worse at their actual job. So which method actually works in practice, pre or post? The study provides a complete blowout victory here. Pre-alignment absolutely crushes post-alignment. Really? By how much? When looking at the data, pre-alignment outperformed post-alignment in over 90 % of the tested tasks.
10:25Over 90 %? That's not even a contest. Why does post-alignment fail so catastrophically by comparison? Well, it comes back to the exact risk you just identified. Post-alignment, especially when dealing with smaller data sets, frequently result in a phenomenon researchers call pseudo-alignment. Pseudo-alignment, meaning the model looks like it understands the numbers, but it's really just faking it. It's essentially severe overfitting. By forcing the LLM to update its internal weights to understand raw numbers, you are actively distorting its incredibly valuable pre-trained parameter. Oh, so it forgets what it already knew.
11:02Yes. It starts memorizing a specific set of numbers rather than actually mapping those numbers to its broader world knowledge. You induce what's called catastrophic forgetting. Catastrophic forgetting. That sounds bad. It is. You degrade the very historian knowledge you brought the LLM in to utilize in the first place. Pre-alignment works because it protects the asset. Right. Keeps the historian's knowledge completely intact and simply provides them with a perfectly translated document. So the key takeaway here is that just plugging an AI directly into your data pipeline is a recipe for disaster.
11:34The method of translation, the pre-alignment, is literally the difference between a 90 % success rate and total failure. It's the critical step. You cannot skip it. But wait, if we assume we have the perfect pre-alignment translator built, let's say that's totally solved, do we actually need to run every single piece of data through a massive computationally expensive LLM? What do you mean? Well, if a business has highly stationary data, meaning the statistical average doesn't really change over time, or highly seasonal data like a company that sells winter coats every November, I mean, isn't using an AI supercomputer massive overkill for that?
12:12Oh, it absolutely is overkill. And the researchers are very clear about this. Okay, good. So it's not a magic fix for everything. Not at all. The study establishes strict capability boundaries. They analyze the statistical properties of A-data sets, things like stationarity, seasonality, and trend, to find exactly where the LLM actually provides an edge. So they found the sweet spots. Yes. They isolated two specific statistical conditions where the LLM proves genuinely indispensable. Let's break those down. When is the supercomputer actually the right tool for the job? The first critical condition is what they call shifting.
12:47Shifting. Okay, what does that look like? This occurs when data distributions change wildly and unpredictably over time. So scenarios where the past is suddenly a terrible predictor of the future. Things like the outbreak of a pandemic, a sudden stock market crash, or highly erratic weather events. That's exactly it. The study highlights that data sets tracking COVID spread, fluctuating exchange rates, and erratic weather have exceptionally high shifting values. And traditional models struggle with that. They completely fail. In these environments, a traditional lightweight model breaks because its historical rules just stopped applying.
13:22But the LLM's pre-trained knowledge thrives here. Because it understands the context. Yes. It has the broader semantic context to understand that a sudden anomaly isn't just a glitch in the spreadsheet. It's a fundamental shift in the environment. It adapts its predictions based on the underlying structure of a crisis. That tracks perfectly with the historian analogy. The historian knows that wars or pandemics rewrite the rules of the economy, whereas the sheltered mathematician just assumes their formula broke and gives up. Precisely. Okay, so that's shifting. What is the second condition? The second condition is high transition.
13:59This refers to frequent complex state changes within the data itself. Imagine multi-stage dynamics where a system is constantly evolving through different phases very quickly. So highly volatile chaotic data with lots of sudden micro changes like analyzing high frequency trading data or maybe a power grid during a massive storm. Yes, perfect examples. And in these high transition environments, it's not just the LLM's world knowledge that provides the advantage, but its actual physical architecture. The way it's built helps it. Right. The sheer computational depth of a transformer model allows it to map out a chaotic web of transitions far better than basic smaller models ever could.
14:40So synthesizing this for anyone managing data right now, if your business tracks stable, predictable trends, your stationary and seasonal data, keep using your traditional lightweight models. Save your computing budget. Absolutely. Don't waste money on an LLM for that. But if you are dealing with chaotic, shifting real-world data where the rules change constantly, the LLM isn't just a luxury buzzword. It's a lifesaver. The findings confirm that entirely. In fact, the study showed that when shifting in the data is weak, satisfactory performance can often be achieved using just the basic encoder and decoder modules of the pre-alignment phase.
15:18Wait, meaning you bypass the LLM completely? Completely bypassing the LLM altogether. Wow. But wait, if the system can just bypass the supercomputer when things are simple, how do we know the LLM is actually doing the heavy lifting during the chaotic scenarios? Ah, I see where you're going. Yeah, how can we be absolutely sure it's not just acting as a bloated, incredibly expensive filter while the smaller pre-alignment network does all the actual math? That is the ultimate skeptics question. And honestly, it leads to the most fascinating mechanical proof in the entire study. OK, let's hear it.
15:49How do they prove it? To prove the LLM's true value, the researchers didn't just build a static pipeline. They utilize a really cool technique called token level routing analysis. Token-level routing analysis. Break that down for us. How does a model route data token by token? Well, they built a dynamic system with a built-in router. For every single token, which in this context is a sequence of time series data, the router calculates the complexity and the entropy of that sequence in real time. The model then literally makes a split-second decision. Do I route this complex data sequence deep through the massive LLM, or is it simple enough that I skip the LLM entirely and process it with a faster, simpler pathway.
16:30Oh, that's brilliant. It's like a traffic cop standing at a fork in the road looking at every single car. Does it send the data down the superhighway or the local road? Exactly like that. So what did the traffic cop decide when they looked at the results? What happened? The results were mathematically undeniable. The analysis proved that when a data sequence exhibits high shifting or high transition, the router deliberately chooses to send the data directly into the LLM. Wow. So the system itself recognizes when a data sequence is too chaotic for basic math, and it actively chooses to invoke the historian to make sense of the chaos.
17:07Yes. But when it sees a nice, straight, stable line, it just routes it around the LLM to save computing power. The correlation is incredibly strong. The decision to route tokens through the LLM directly correlated with minimizing forecasting errors in those really complex scenarios. That's amazing. It provides It's mechanistic mathematical proof that the LLM isn't just a passive filter. It is actively saving the day when the data gets tough. It's an adaptive utilization of the LLM's vast capabilities. That routing mechanism completely shuts down the argument that the LLM is just dead weight. But, you know, I can already hear the executives out there thinking, great, complex data needs an LLM.
17:45I have complex data. So I'll just go buy the biggest, most expensive model with the most parameters on the market, throw it at my chaotic data, and I'll get perfect predictions. Yeah, that thought process leads straight into what we call the magic bullet fallacy, and the study explicitly warns against it. So bigger isn't always better. Scaling up the model size does not automatically yield better forecasting. More parameters do not equal more accuracy if the setup is wrong. Because if the pre-alignment translator is bad, handing a garbled report to a smarter CEO just means they get confused faster.
18:19That's perfectly stated. But to build on that, it's not just about the translator. The study found that under large-scale mixed data distributions, partial fine-tuning is no longer sufficient. Meaning you can't just tweak the model a little bit. Right. You need a fully intact LLM. If you compromise its pre-trained weights in any way, performance drops. You need the historian fully awake with all their vast contextual knowledge completely intact. And you still have to tell them what they are looking at, right? I mean, you can't just slide the perfectly translated numbers under the door without a memo.
18:52That is exactly where semantic guidance comes in. The researchers prove that prompting the LLM with text, actually describing the background information and task specifications alongside the numbers, consistently improved performance. So you literally have to provide a prompt saying, hey, these numbers represent a sudden spike in server traffic during a major holiday sale, rather than just throwing the numbers into the void and hoping the model guesses the context. Yes, the model needs to know which part of its world knowledge to activate. Combining intact pre-trained weights, rigorous pre-alignment, and clear semantic guidance is really the only way to successfully scale these models for forecasting.
19:30Man, we have covered a massive amount of ground today. Let's recap the core mechanics for everyone listening. We started with a very valid skepticism. LLM shouldn't be able to forecast numbers. Right, but we learned that the skeptics were working with limited data, shallow layers, and flawed alignment. Exactly. When evaluated comprehensively across 8 billion observations, the evidence is clear. LLMs definitively improve forecasting, but only under specific architectural constraints. Constraint number one being that pre-alignment is non-negotiable. You have to translate the numbers into the LLM's structural language before feeding it in.
20:10Post-alignment destroys the model's knowledge through overfitting. And constraint number two, apply them to the right statistical environments. LLMs excel in chaotic, shifting, high-transition data. If your data is highly stable and predictable, just stick to traditional models. And finally, size isn't everything. You must use the LLM intact, protect its pre-trained world knowledge, and guide it with clear semantic text prompts to give it context. So if you were sitting in a strategy meeting next week, and someone enthusiastically suggests, hey, let's just throw an LLM at our supply chain data to see what happens, you don't just nod along anymore.
20:46No, you have the tools to push back. You know exactly what mechanisms to question. You ask, is our data genuinely shifting or is it stable? Are we actively pre-aligning the modalities to protect the model's weights? And are we providing the semantic context the AI needs? Or are we just hoping it magically understands an empty spreadsheet? Asking those questions ensures you are actually architecting a real solution rather than just chasing the latest technological hype. And that leaves us with a lingering thought to walk away with today. We started this deep dive, wondering how an AI trained on poetry and Reddit threads could possibly read the cold mathematical numbers of a spreadsheet.
21:25Yeah, it seemed impossible. But if LLMs can use their structural language knowledge to successfully decode the chaotic, purely mathematical shifts of the stock market or global weather systems? Well, what other seemingly non-linguistic fields might secretly be governed by patterns an LLM could read like a book? If the universe is ultimately written in mathematics, maybe a language model is finally learning how to translate it back to us.
From the publisher
This research paper evaluates the efficacy of **Large Language Models (LLMs)** in the field of **time series forecasting (TSF)** through a massive empirical study. While previous scholars argued that LLMs offer minimal benefits over standard models, this study utilizes **8 billion observations** to prove that LLMs significantly enhance **cross-domain generalization** and predictive accuracy. The authors identify that **pre-alignment strategies**, which map numerical data to word embeddings, generally outperform post-alignment fine-tuning. Their analysis reveals that LLMs are particularly powerful when dealing with **distribution shifts** and **complex temporal dynamics** rather than simple seasonal patterns. Furthermore, the paper introduces a **routing mechanism** to show that models adaptively choose when to utilize LLM logic based on data complexity. Ultimately, the findings provide a framework for using **pretrained world knowledge** to improve forecasting across diverse real-world scenarios.




