In short
Podcast Summary: The TWIML AI Podcast - #685 Chronos: Learning the Language of Time Series
Episode Overview In this episode, Sam Charrington hosts Abdul Fatir Ansari, a machine learning scientist at AWS AI Labs, to discuss his paper "Chronos: Learning the Language of Time Series." The duo explores the challenges of applying pre-trained language models for time series forecasting, the advantages of the Chronos model, and its promising results in zero-shot forecasting benchmarks.
Key Points Discussed
Guest Background
- Abdul Fatir Ansari:
- Transitioned from a bachelor's in civil engineering to machine learning, earning a PhD in deep generative models.
- Developed an interest in time series forecasting during an internship at Amazon, leading to a full-time position focusing on this area.
Introduction to Chronos
- Inspiration for Chronos:
- Observed that many production time series systems use simple models due to the complexity of implementing state-of-the-art models.
- Simple models often perform surprisingly well, leading to hesitance in adopting more complex models.
Time Series Problems
- Common applications:
- Demand forecasting in various domains (e.g., energy, retail).
- Critical for labor planning and operational efficiency.
Traditional Time Series Models
- Two Broad Categories:
- Statistical Local Models (e.g., ETS, ARIMA):
- Fit a unique model per time series.
- Effective for limited data and low-frequency time series.
- Task-Specific Deep Learning Models (e.g., DeepAR):
- Train a single model across related time series.
- More flexible and effective with larger datasets.
Challenges in Time Series Forecasting
- Overfitting:
- More prevalent in deep learning models, particularly when limited data is available.
- Benchmarking Issues:
- Reliance on a narrow set of datasets can lead to misleading results.
Chronos Model Overview
- Innovative Approach:
- Utilizes a language model architecture but trained from scratch on time series data.
- Employs a unique tokenization scheme to convert time series data to a discrete signal for training.
Tokenization Scheme
- Normalizes time series data and quantizes it into bins, which allows for using existing NLP architectures without modification.
- This design choice aims to simplify the modeling process while still achieving effective results.
Performance and Evaluation
- Zero-shot Performance:
- Chronos performs comparably to task-specific models on unseen time series data.
- Demonstrates better performance than traditional statistical models in certain contexts.
Data Augmentation Techniques
- Mixup and Gaussian Process:
- Combines real-world time series data to diversify patterns during training.
- Uses Gaussian processes to generate synthetic data, helping to improve model robustness.
Future Directions
- Improve the quantization scheme to better handle sporadic spikes in data.
- Enhance the quality of synthetic data to further bridge the gap between real and synthetic datasets.
- Explore advanced modeling techniques from recent language models to improve performance.
Critiques and Responses
- Acknowledgment of critiques concerning performance compared to statistical models.
- Emphasis on the importance of open dialogue and continuous improvement in machine learning research.
Conclusion The conversation between Sam and Abdul highlights the complexities and innovative approaches to time series forecasting with models like Chronos. The discussion emphasizes the importance of incorporating machine learning advancements into operational models while recognizing the challenges presented by traditional forecasting techniques.
For further details and complete show notes, visit [TWIML AI Podcast - Episode 685](https://twimlai.com/go/685).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00This is also one of the problems in the literature. In many papers, they are using the same set of datasets. And at some point, improvements on those benchmarks don't really mean anything because there is no real absolute improvement in performance over time.
0:25All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington, and today I'm joined by Abdul Fethir Ansari. Abdul is a machine learning scientist at AWS AI Labs in Berlin. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Abdul, welcome to the podcast. Thanks a lot. Thanks a lot, Sam, for inviting me. I'm looking forward to digging into our conversation. This time around, we'll be talking about your work in modeling time series, and in particular, the Kronos paper. Before we jump into that, I'd love to have you share a little bit about your background and how you came to work in the field.
1:05Yeah, sure. Thanks a lot. My background is actually a bit interesting because my bachelor's was in civil engineering. It was not really computer science or machine learning. But after that, I moved to Singapore where I got my PhD from NUS. And there I worked mostly on deep generative models. And that's the field I'm broadly interested in, everything related to generative models, which basically means modeling, representation, learning, inference, these kinds of things. And for time series forecasting, I actually did an internship at the same team at Amazon. That's where I got interested and I got introduced to time series modeling.
1:38And since then, I've been working on time series forecasting. And then I joined Amazon full time, the same team, and I've been working on time series since then. Awesome. Awesome. Was there a particular event in your academic experience that was the bridge between civil engineering and machine learning? Yeah, I wouldn't say so. I think I was always interested in computer science and machine learning since the beginning. It's just a weird sequence of events that led me to get into civil engineering. Okay, okay. So let's talk a little bit about the paper. What was the inspiration for taking on this particular project?
2:15I think the major inspiration actually came from the things that we see in production time series systems, because typically people use very simple models in production systems. And the reason for that is there is a lot of work that needs to be done to put in a new state of the art model into any production system, because first you need to set up some kind of data pipeline to collect data, then train models on this data and then do a lot of evaluation before you can do some reasonable change of models. And this is a lot of steps and people tend to kind of not do all of this stuff and use very simple models.
2:52And time series forecasting is a bit special in the sense that where very simple models, in fact, something like seasonal naive, which basically takes the previous season as the forecast, seems to do very well for many kinds of time series data set. So people tend to choose and use these simpler models compared to state-of-the-art deep learning models. And there's also this kind of hesitancy amongst people to not use state-of-the-art models because in the academic domain, you have a few data sets that people kind of overfit on. And it's very easy to get misled by those results and then hope that that model actually works on real-world time series, which often is not the case.
3:30So you referenced real-world time series. What are some of the types of problems that you are trying to apply real-world time series analysis to? There are actually a lot of problems both inside Amazon and outside where people use time series forecasting. Like most importantly, it's used for some kind of demand forecasting, like, for example, for energy demand forecasting or also for retail forecasting. I mean, Amazon is one of the biggest e-commerce companies. So there's also this retail forecasting component, which is also not unique to Amazon, but in general for labor planning, you need to do forecasting and many other things.
4:07So there are actually a lot of problems that use time series forecasting, at least in the loop, not really as an end goal. So you would want to use your forecast somehow to take some kind of action. So it's not really just to get a forecast, but to actually use that forecast somehow in a downstream system. And you mentioned some of the traditional approaches to time series. Can you provide kind of an overview of how the traditional statistical approaches tend to work? Sure. So I think for the existing models for time series forecasting, I would broadly classify into two categories. But of course, there are many other different kinds of classifications.
4:47One of them are so-called statistical local models. And the idea for these models is that you fit one model for each individual time series. So you can just look at one time series. So models like ETS, ARIMA, these models basically fit one model or one system for each time series. The second kind of models, I would say, are task-specific deep learning models, like DeepAR, PESH-TSD, or all the recent models that are coming out. I would call them task-specific deep learning models because they look at a bunch of time series that are somehow related. And then they train one model, so-called global model, across all of these time series.
5:26And then you can use this model for inference on some downstream time series. So these are the two broad categories. Of course, there are positives and negatives for both of these techniques. for simple statistical models I would say the positives are that they are extremely strong baselines like I mentioned previously especially in the case where you have limited data they tend to do very well and also for low frequency time series like monthly, yearly weekly this kind of data they seem to do very well because typically there you don't really have a lot of data but as soon as you get a lot of data deep learning models tend to do better because then you can train your model on a large corpus of such time series and then you can hope to get better improvements.
6:08There's another point of difference between these two models in terms of flexibility. I would say the first model, the local models, they train one model for each time series, which makes them slow and less flexible. And also they have some kind of design assumptions for how time series should be decomposed, what kind of error distributions you are modeling. But in the case of deep learning models, of course, all bets are off, right? You can really have any kind of arbitrary design assumption in a deep learning model. So in that sense, they are more flexible and they can also be used for multiple time series together.
6:43So these are some positives and negatives. Of course, more recently, you have these pre-trained models, just like Kronos. The idea there is to train one model on a large corpus of time series data and then hope that on an unseen downstream task or downstream time series, it would do well. And that's what we see, at least in our work. You mentioned overfitting as something that you see a lot of with, is that applied to traditional models in particular or the deep learning models that you mentioned? I would say overfitting is probably more of a problem for deep learning models. when you have limited data and a large-ish model, you can tend to overfit very easily.
7:28And this is what we see typically. And this is also one of the problems in the literature. Not necessarily any specific paper, but in many papers, they are using the same set of data sets. And at some point, improvements on those benchmarks don't really mean anything because there is no real absolute improvement in performance over time. And the field kind of is converging to some performance. There's always some delta that you can show compared to some selected baselines. But in terms of absolute performance, things don't seem to be improving. So we definitely need better data sets, better benchmarks for sure to say that this model is better than some other baseline model.
8:14Got it. Got it. And so the approach that you took with the Kronos work is to try to leverage language models. Is this the first paper that tries to leverage LLMs, or in what ways have folks tried to leverage LLMs to solve the same type of problem? This is definitely not the first work. I would like to say that this is one of the good attempts, but this is definitely not the first work. In fact, if you look at Neodips last year, there were several papers that are trying to leverage language models somehow. So I would broadly classify them into two categories. Some models that just try to use pre-trained language models directly to prompt them in a certain way for time series forecasting.
8:58So one of the good examples of such a model or such a method would be LLM time, which was in Neodips last year. It basically uses a LAMA 70 billion or GPT-3 model for basically doing forecasting. you convert your time series somehow into a string of numbers, and then you just prompt the model using this string. The second category of model that uses LLMs for forecasting would be like time LLM or GPT-40S. So these models take some backbone language model architecture like GPT-2, and then they have some kind of fine-tuning scheme that you can adapt the model weights for time series forecasting. So these are the two main categories that use language models for forecasting.
9:39But even beyond that, there are other works like, for example, Moirai from Salesforce and also Laglama and a few other works, TimeGPT from Nixla. These models basically, they train a transformer-based model on a large corpus of time series data. They're not really language models per se. They're just like large transformer models trained on a lot of time series data. So in some sense, Kronos is somewhere in between these kinds of works. So it uses a language model architecture, but it also trained from scratch on a lot of time series data. So it's not really using a pre-trained language model like some of the previous works.
10:18Got it. Got it. Got it. And it seems like one of the core ideas around the way that you've attacked this problem is the tokenization scheme. Yeah, exactly. One of the key motivations of our work was to be as simple slash as lazy as possible. So the idea is, why reinvent the wheel? Why build something new? We just wanted to verify what the things, the infrastructure, the architectures, the libraries, everything that's already there for natural language processing. Can we directly use it for time-size forecasting? In fact, in this work, we don't make any changes to the language model architecture.
10:55And it's also surprising that we also don't make any changes to the objective function, right? for language modeling, the objective function that's used typically is cross-entropy. So it's a classification-based loss. And it might be counterintuitive to say that for time series forecasting, we are using cross-entropy or a classification loss. So essentially, we are doing regression via classification. And it tends to do really well, which is surprising and not surprising because there are also some existing works that already do this. For example, if you are familiar with WaveNet, it was a generative model for audio.
11:28And audio signals are also time series. They also have a similar kind of approach to basically use like bins for modeling the audio signal. And we are also using a similar kind of approach. So at the core of it, it's really simple. So we don't have any time series specific design choices or design modifications from the architecture perspective. And also from the loss function perspective. We just encode the time series somehow into a discrete signal from a fixed vocabulary. And then we train the language model on a lot of such tokenized time series data. And what does that mean or look like to tokenize a time series data set or to map it to a fixed vocabulary?
12:16Yeah. So, of course, tokenization is a broad general term. but the specific scheme that we are using in our work is we take the original time series which can be at any arbitrary level so it can be really millions of something or between zero and one so it can be at an arbitrary scale so the first step is to somehow normalize the time series into a reasonable range and this normalization can be can be a design choice in our work we use mean scaling which just scales the time series by its absolute mean and then what what that does is it squishes the time series down into some reasonable range.
12:52And once it comes into a reasonable range, we select some bounds on the real line and we basically construct uniformly spaced bins between these bounds. And we take the scale time series and we quantize them into these bounds. And after quantization, you basically get the indices of these bins, which also kind of looks like the original time series. But of course, there is some loss of precision here because the bins have non-zero width. So there is some loss of precision, but you have somehow converted your original continuous signal into a discrete signal from a fixed vocabulary. And then once you have done this kind of tokenization, you can essentially train any kind of language model architecture on this.
13:33So it doesn't really matter which architecture you select. So in the paper, we really focused on the T5 encoder-decoder architecture, which seems to do very well. But it's not really a restriction. In fact, you can choose an architecture which is not even transformer-based. So, for example, recent works that do state-based modeling for lateral language processing like Mamba, you could also use such kinds of architectures. So it's not really a restrictive framework. Is there a simpler possible way that you could have done the tokenization? That seems like a very straightforward approach to it. I'm not sure if you can be simpler than that.
14:11But if you have ideas, I'm open to discussing them. But I feel like this is the simplest possible. In fact, just for quantization, in prior literature, there are some other techniques that one could use. So, for example, here you have uniformly spaced bins. But what you could do is something called data-dependent binning. Something like quantile binning, where basically you select your bin edges such that equal number of points fall in each bin. But the reason we didn't select such kind of approach is then this approach really depends on the training corpus. And in the downstream inference time, you could have such data which probably was never seen during training and the distribution would be completely different.
14:54And you might see some worse performance on downstream data. So that's the reason we went with the simplest possible thing. But I'm not sure if you can go even simpler. It's extremely simple. That's what really struck me. Like I can imagine ways to get arbitrarily more complex, but that seems like the most basic place to start. And, you know, as it often happens, it surprises me that it works. Yeah, it's definitely very surprising. And so you mentioned that you use T5, but you didn't need to use T5. Is there a reason why you chose T5 for the paper? And were there other language models that you tried?
15:35I have a very boring answer for it. The boring answer is that I just started with T5. So when we did our initial experiments, we just started with T5. And the reason for starting with T5 is because T5 is extremely popular in the NLP community. I mean, these decoder-only models are only getting popular very recently, but T5 has been popular since its beginning. The second reason is all of these recent decoder-only models, their configurations and everything were typically available in very large sizes. But T5 comes from a tiny size, which has only 8 million parameters, to an Excel size, which has a few billion parameters.
16:13So there was a wide range of configurations, and we knew that these configurations work very well for NLP. So that's the main reason that we selected the T5 because it came in a lot of different configurations that people already knew works well for NLP. So we didn't want to redesign or reinvent the wheel. That's the main reason why we selected T5. Got it. Got it. Another element that the paper calls out as being important to the work is the use of generated data or data augmentation. Can you talk a little bit about how you incorporated that? Sure. the modeling framework or any kind of loss function, whatever, architecture design, I think it's only one side of the story.
16:55And as more and more papers in language modeling are also saying that data is the key, right? Data is the solution to real good quality language models. I would say same is the case for time series forecasting. And one of the special limitations for time series forecasting is that you don't actually have so much good quality publicly available time series data. Many companies have some internal data sets that probably they won't release that are probably much more high quality. But in public domain, we don't have such high quality data sets. And you say that, but you also mentioned in the paper that you tested against 48 data sets, which seems like a big number of data sets.
17:36Yeah, exactly. It seems like a big, it seems like a, so by the way, we tested on 42, but it seems like a big number. but some of those data sets may have only a few time series, like 10. So at the end of the day, in total, we collected 890 ,000 time series. But there is probably an over-representation of certain domains and a very limited or under-representation of some other domains. So this data set, I think comparing to the language modeling data sets, which is essentially at the scale of the entire internet, this is significantly less. When you say 890 ,000 time series, I guess the other unknown variable is the length of those time series.
18:19Exactly. To compare this to other approaches to training language models, does it make sense to boil it down into a number of tokens that were used in training? Do you know that number? If I remember correctly, it's something around 84 billion. But I would say that for language modeling, the tokens are much more high quality and diverse compared to time series forecasting. Because here, for example, there is an over-representation of Wikipedia page visits data set or weather bench, which is weather data set, right? So it might have very similar characteristics for multiple cities, and it may not bring so much additional information compared to the case of language modeling.
19:01And that's the main reason why we had to try these augmentation schemes to basically improve the performance even further. Got it. So just to make sure I understand that in the case of language modeling, each additional web page or book or aspect of data is kind of fairly diverse. Whereas it sounds like you're saying that many of the time series in your data set are essentially different slices of the same generating function. So, you know, visits to, you know, the same website, that kind of thing. Yeah, I mean, exactly. I don't really have data to back this up, but based on the sizes of the data sets.
19:44So for example, the weather events data set that we used was fairly large. The Wikipedia page, which is data set that we used, is also fairly large. So it contributes a lot to this 890 ,000 time series. And the diversity, I would say, comparing to the language modeling data sets is low, but I don't have specific data points to back this up. Fair enough. So you were talking about incorporating data augmentation. Yeah. So in particular, we looked at two augmentation schemes. One of them essentially just takes random real-world time series from different data sets. They may or may not be from different data sets.
20:23And then just takes convex combinations of these time series. And the idea here is to just diversify the types of patterns that the model sees during training. So, for example, from some data set, maybe you see some kind of trend. From some other data set, you see some other interesting seasonal patterns. If you take a convex combination of these two data sets, you will get a time series with combinations of these two types of patterns. And that's essentially what this augmentation scheme does. We also have an example in the paper showing what exactly it does. The second scheme that we used is to generate new synthetic data using Gaussian processes.
21:00So it's an extremely simple scheme. the core idea is to have some kind of bank of Gaussian process kernels. And so these kernels would represent some base functions or base time series. So for example, from a linear kernel, if you draw a signal, you get a linear function. From a periodic kernel, you get some kind of seasonal time series. And the idea is to randomly draw these kernels from this bank and combine it using additional multiplication, which also gives you a valid kernel. And once you have these more complex kernels, you can draw samples from these kernels. So let's say you have two linear kernels and you take their multiplication.
21:35So you get like a quadratic kernel, which would give you some kind of quadratic trend. But then if you add it to a periodic kernel, you get some quadratic trend with seasonal properties. So you can basically construct very complex time series using this very simple scheme. And in the paper, we show that these augmentation schemes are extremely helpful, especially in the case of zero shot performance, which is what really is the thing that we care about at the end of the day. So on an unseen time series, how does the model perform? And adding such kind of augmentation schemes and synthetic data really improves the performance of the model on unseen time series tasks.
22:13And so what percentage of the ultimate training data set was synthetic versus the data sets that you sourced? So we have a 90-10 split. So 90 % of data comes from the real augmentations and 10 % of data comes from synthetic data set. So this was the original scheme that we used for the main models. But then we also did some ablation, which shows that this is a near optimal choice. So if you select 90 % real and 10 % synthetic, you get almost the best performing model on all metrics. But we also trained with different percentages. For example, we also trained a model that was purely trained on synthetic data, just to see what we can gain on only training the model from synthetic data.
22:58And interestingly, this model also doesn't do very bad. So it's actually one of the models that does reasonably well. It's better than some baselines, which is very, very interesting. Of course, it's not as good as the model trained with both real and synthetic data, but there could be potentially ways to improve the quality of synthetic data to reduce this gap between real and synthetic data. And this, I would say, is slightly easier compared to the case of language modeling, because synthetic data in the case of language modeling really means you first train a big model, and then from this big model, you generate synthetic data.
23:35But here, there's a very simple scheme. You don't, so it's not a chicken and egg problem here, because here you don't really need data a priori. You can come up with a scheme like using Gaussian processes, like we did in the paper, to draw reasonably looking time series data. Of course, what data, what time series is good or bad, this problem is really unsolved, at least at this point, if this time series would really help in the downstream performance. But that's an interesting way to augment using purely synthetic data sets. Okay. I thought I heard you mentioned, were you distinguishing between the real data sets that has been augmented using the combinations and synthetic?
24:22When you talk about the 90-10 split, is the 90 strictly real data sets and the 10 is a combination of the TS Mixup and KernelSynth or was TS Mixup stuff included in the 90? So the 90 is actually purely TS mixup. And because in TS mixup, you have a parameter k, which basically tells you how many time series you want to draw. And this k can range from one up to some value k. So if you sample k equals to one, you get the original time series. So original data is also represented. But this 90 % includes all the augmentations, possibly with original time series, but also with combinations. Got it, got it.
25:03So you're distinguishing real and augmented data via TS MixUp from the synthetic data via KernelSynth, which is the Gaussian process stuff. Yeah. The 10 % is only KernelSynth. Got it. Okay. Awesome. And then so how do you, you evaluated against a series of historically relevant academic benchmarks? Yes. So we used many data sets from, for example, Monash benchmark, which has a lot of time series data sets, but also some other data sets that are popular, like electricity data set, ETT data set. There's several other data sets that are popular in the literature. Our specific benchmarking or evaluation setup is different from these existing bugs.
25:51Like, for example, there's a long-term time-series forecasting benchmark that's quite popular in many papers. But we strictly went with a new benchmark setup so as to not give any model any kind of advantage. So we just use these data sets, but the evaluation setup is different and it's the same across all the models that we're using. Okay, can you talk a little bit about the evaluation setup then? Yeah. So for the evaluation setup, basically we just selected, so for each kind of dataset based on its frequency, we selected some splitting points so that not to leak any data for any model. And then basically with some prediction length, that's it.
Read the full transcript
26:33So this may be same. So for example, for some benchmarks that come from competitions like M competitions, these kinds of benchmark, we kept the setup same as it was used in the original competition. But for some other data sets, the setup is slightly different, basically potentially with different prediction lengths. So it may or may not be the same with some existing works. But it was uniform across all the models. So there was no advantage given to any specific model for this. And maybe going back to our earlier conversation about the risk of overfitting, do you see any risks that this model has previously seen aspects of these data sets and its training data?
27:16It's possible that there are some patterns that the model has seen, maybe from kernel synth or some augmentations. But like specifically from the zero-shot test set, there is no time series that was directly in the training corpus. But in terms of patterns, potentially the model has seen maybe due to synthetic data generation or some kind of augmentations. but directly the time series were not part of the rendering corpus, at least for the zero-shot datasets. And can you talk a little bit about the, ultimately what you're trying to do here is create a model that can do zero-shot forecasting. Can you talk a little bit about the end result, how well it performed in zero-shot?
27:59So especially in the case of zero-shot benchmark, we obtained very promising results. the highlight of the benchmark is that the performance of zero-shot Kronos models especially the best performing large model is on par with task-specific models that were trained on those data sets so Kronos is completely zero-shot here it has never seen these data sets during its training but these baseline models were trained on these data sets individually so the performance just the fact that the performance is close to each other tells you that Kronos is a very promising model. Of course, the Kronos is not the best performing model.
28:38So for example, if I remember correctly, in the probabilistic benchmark, DFT does better than Kronos, but the performance is fairly close. So the zero-shot performance is fairly close to this DFT model, which was trained on these data sets. So in terms of performance, the gap is not that large. Another thing that I would like to call out is when you compare Kronos with some other models that are out there for pre-trained time series modeling. So for example, Moirai models from Salesforce, Laglama, and LLM time, a few other pre-trained models. We see that Kronos actually performs much, much better than these baseline models.
29:17Especially for example, if you compare Kronos with Moirai, on the zero-shot benchmark, Kronos is much better, even though Moirai may potentially have seen some of these zero-shot data sets because their training corpus and our training corpus was not really the same. So they might have seen some of these data sets. The model may have seen some of these data sets, but even then, Kronos' zero-shot performance is better than these pre-trained models. Another thing to highlight is local statistical models such as ARIMA and ETS. These models are typically used in an inference-only setting because they don't really require any training on a large corpus.
29:54So you can give it a new time series and they will basically just fit a model for this time series. so Kronos has a similar kind of setup so you can give it an arbitrary new time series it will do forecasting for you but when you compare Kronos with these baseline Auto ARIMA, Auto ETS these kinds of models Kronos is much better and at the same time it's extremely fast compared to these models because for example for the Auto ARIMA model you would train a lot of models and then you would based on some criteria select the best model which takes a lot of time especially for long time series with high frequency then that's the reason Kronos is much faster than these kinds of simple baselines like auto-remin, auto-DS, while being much better in terms of the performance.
30:40Okay, let me try to summarize that. There are, generally speaking, two ways that you would use this model for forecasting. The first is zero-shot, which is essentially trying to use it the way we use language models. You have this pre-trained model. You've never seen a particular time series and you give it some segment of the time series and you want it to project into the future. That's kind of the zero shot scenario. And then you've got what you call in the paper in-domain. Is in-domain versus zero shot, is it only a question of benchmarking? And what I mean by that is when you're using Kronos, are you using it any differently in domain versus ZeroShot?
31:27You're still just giving it some data and it's projecting. So from the Kronos perspective, it doesn't really matter. It's just a question of if you're comparing against an autoregressive model versus something that is ZeroShot. Exactly. So in domain and ZeroShot, this is just like a synthetic classification. it doesn't really matter for Kronos if this time series was really in the training corpus. It just matters from an evaluation perspective because if you show extremely good performance on in-domain, one might claim that all of these time series were part of the training corpus. So that's why the zero-shot performance goes hand-in-hand, which is probably the more interesting case, right?
32:07So you won't pre-train your model from scratch every day, for example. But in terms of zero-shot, you can use it on any arbitrary time series data set. Yeah. Yeah. And so your point then about time that it takes for a traditional model versus the using Kronos is that with the traditional model, you have to train that traditional model and you're incorporating that time into, I guess, the way you're benchmarking as opposed to Kronos, which is essentially zero shot. Exactly. In both cases. So for the traditional models, there is a plot in the paper which shows the inference time. For traditional models like Auto-OTS and Auto-REMA, there is no real distinction between the training phase and the inference phase.
32:55So all the times gets accounted for in the inference time. But for the deep learning models, which have a clear distinction between a training phase and the evaluation phase, we only compare the inference time. So in terms of inference time, these deep learning models would be much better than any pre-trained models because they are typically smaller, but they also have the training time, which is not accounted for in the plot that we have in the paper. So the training time, I mean, training also takes a lot of time. And if you account for both times, then the zero shot inference time of Kronos is much better because it doesn't really need to do any kind of training.
33:29You can just feed it the time series and get the forecast. Got it, got it. And did you find any particular types of patterns that the model performs better or worse for, you know, meaning seasonality or other types of trends? Or is it, you know, fairly agnostic to, you know, types of time series patterns? So in the paper, we actually have a reasonably sized section discussing all of these properties of the model. and one thing especially that I would like to call out where the model tends to not do so well is sparse data when you have spikes. So, and this is essentially the problem. It's not really a problem of the complete framework, but it's a problem of the tokenization scheme that we're using currently because we are doing mean scaling and if most of the time series is zeros, but then you suddenly have a few spikes, even if the spikes have some regular pattern, these spikes may not be represented well due to loss of precision.
34:27So when the spikes are not represented in your input, the model, of course, won't be able to forecast them correctly. We show that the sparsity of the spikes actually does matter. So if you have frequently occurring spikes, then you can represent them well. But if you have sparse spikes, then the tokenization scheme may not be able to represent these spikes and the model may not be able to forecast them. So this is one of the cases where it doesn't do so well. But we have shown many other cases, especially qualitatively, that, for example, for seasonal data, many kinds of seasonal patterns, the model does very well, both in the in-domain and especially in the zero-shot setting.
35:05On completely new data sets with different kinds of seasonality patterns, the model tends to do very well in catching these seasonality patterns. But when you have spiky data, the model tends to struggle, at least the current models that we have, because they use a specific tokenization scheme, which may not be able to represent these spikes. and that kind of gets you back to the what you discussed earlier in terms of not necessarily being stuck with this quantization approach you could iterate that or make that more complex to deal with some of these issues like spikiness exactly so this is this is i would say one of the most interesting next steps to consider especially things where time series where the model is failing where this current tokenization scheme is not doing so well.
35:51So there are some quick inference time fixes. So for example, if due to scaling, the tokenization scheme doesn't represent your data, maybe you can experiment with a different kind of scaling scheme as a preprocessing step. So for example, doing standardization or some kind of transformation, but these are kind of inference times hacks. What would be amazing to have is some kind of tokenization scheme that really doesn't struggle when you have such kind of sparse, spiky data. But this is really the next step and not part of the Kronos work at the moment. I came across a tweet from the Nishla folks that kind of criticized this.
36:29You could say, I guess their finding was that it was 10 % less accurate and 500 % slower than classical statistical models? Have you come across that critique? And what do you think about that generally? Yeah, yeah. Thanks for bringing this up. Yeah, we have come across this critique. And in fact, one of the main reasons for open sourcing these models is that people can use them, they can break them, criticize them in all the possible ways. And that's how we grow as scientists and that's how science moves forward. We came across this critique. And in fact, we have also given a response to this critique.
37:08And the conclusion is that the original benchmark from Nixcla only uses a few data sets from our zero-shot benchmark. But if you expand the same evaluation to the complete benchmark, we see a different picture. We see that first, they're comparing Kronos models against an ensemble model of four different statistical models. So this model is actually very strong as a baseline. When you compare the zero-shot performance of Kronos against this trained model, we see that Kronos performs on par with these statistical models, but it's also significantly faster on average across these datasets. So on some selected datasets, it may be slower because these datasets are typically low frequency time series, which are very short.
37:51So these statistical models tends to be very fast on those datasets. But if you look at the complete benchmark, you see a different picture. But in general, we really welcome these kinds of critiques. And there are also some other folks that have shown the limitations of Kronos. And we are looking into that. We are trying to solve these problems so that we can have better and better models moving forward. So it sounds like one key idea here is that Kronos will perform better relative to traditional models as the time series sample that you're trying to predict from gets longer because of the training time aspect that you mentioned previously.
38:35Is that the right takeaway? It's not really about the context length, but basically the kind of benchmark. So one thing, for example, that you observe is that for higher frequency data, for example, hourly, half hourly, 15 minute, this kind of data, Kronos tends to do better than the traditional models. Because for low frequency data, like quarterly, yearly, the time series are typically not very long and you have fewer time series. So potentially those patterns are, in a zero-shot sense, are difficult for a pre-trained model to forecast well. But for a statistical model that is fit on these patterns, it tends to do well compared to a zero-shot model.
39:18That's my takeaway, at least from the complete benchmark. Got it. Got it. And so the approach that they compared Kronos to is an ensemble one, like you mentioned. Do you see any merit to including something like Kronos into an ensemble that also includes traditional types of approaches? Yeah, that's a great point. In fact, we have recently integrated Kronos into Autoglone, which is an AutoML framework for different kinds of data, especially tabular data, but also for time series. And there are ways of ensembling Kronos models with traditional techniques like tree-based techniques or some other statistical models.
39:57And we see that when we include Kronos in these ensembles, you get much better performance compared to just using the base ensemble. So there is definitely something that Kronos is rigging to the table in terms of performance in the end. To what degree is Kronos used in production in Amazon? Is it strictly a research effort at this point or has it started to be incorporated into real world systems? So I can't really answer this question really due to some restrictions. But there is definitely a lot of hope and a lot of people that are experimenting with Kronos, both internally and externally at Amazon.
40:41And hopefully it will get integrated. So for now, the information that's out there in public, it's already integrated into autogluon. And autogluon is used by a lot of folks, both internally and externally. but we are hoping to bring Kronos into other platforms as well. Awesome. Awesome. You mentioned that experimenting with the quantization scheme is a key area of future research here. Are there other standout areas that you are looking to iterate on? There are several areas. One of them is, of course, about synthetic data, how you can improve the quality of synthetic data to bridge the gap between real and synthetic data.
41:22there are some areas of course in terms of modeling because d5 is kind of an old architecture can we can we basically use these ideas from more recent language modeling architectures to to model better there are of course some directions in terms of inference quality and inference speed many things you could directly borrow from the nlp literature but many things you you probably would have to design for for time series forecasting to just use these written models, but with better inference speed and better inference quality. So these are some of the directions that I think are the most interesting at the moment moving forward from Kronos.
41:59Awesome. Awesome. Well, Abdul, thanks so much for joining to share a bit about Kronos and the way you've approached time series forecasting. Yeah. Thanks a lot for inviting me. It's been a pleasure talking to you. Thank you.
42:21Thank you.
From the publisher
Today we're joined by Abdul Fatir Ansari, a machine learning scientist at AWS AI Labs in Berlin, to discuss his paper, "Chronos: Learning the Language of Time Series." Fatir explains the challenges of leveraging pre-trained language models for time series forecasting. We explore the advantages of Chronos over statistical models, as well as its promising results in zero-shot forecasting benchmarks. Finally, we address critiques of Chronos, the ongoing research to improve synthetic data quality, and the potential for integrating Chronos into production systems.
The complete show notes for this episode can be found at twimlai.com/go/685.




