9 ms·
Moirai: A time series foundation model for universal forecasting
- gorold 3y agoCode is available! https://github.com/SalesforceAIResearch/uni2ts https://github.com/SalesforceAIResearch/uni2ts
- peter_l_downs 3y agoAnd the documentation makes me think they did a great job making this easy to use. Looking forward to playing around with it. Edit: oh you’re one of the authors — thank you, and congratulations!
- fbdab103 3y agoChoosing beggers and all that, but the LOTSA dataset could really benefit from a Readme on HuggingFace. Even just a citation back to the original paper would be good.
- gorold 3y agoThat's actually a great suggestion, thanks! We're also still working on improving the readability/usability of the codebase too
- lngnmn2 3y ago[dead]
- mikeyouse 3y agoLooks super interesting. Definitely going to play with this though it took me way too long to figure out what Salesforce Air Search was.. maybe that's a sign I should log off for the day.
- dcl 3y agoWhat does 'any-variate' forecasting mean? Can you use this pre-trained model to produce forecasts when there exists useful covariates/features/predictors? Is this something the other TS foundation models can/cannot do?
- gorold 3y agoWhen we deal with many different multivariate time series, each time series can have a different number of variates. So "any-variate" means that the model is able to take as inputs multivariate time series with arbitrary number of variates, and model the interactions with the Transformer's attention mechanism. This is something that many other TS foundation models do not consider yet - they convert all multivariate time series into multiple univariate time series. Whether or not the forecasts improves as a result of the additional covariates is still an open question which needs to be studied more -- we need to build better evaluations and benchmarks for this.
- dcl 3y agoUnderstood, thank you. There are certainly applications in demand sensing/demand forecasting where things like recent order information, recent sales, CRM inputs are quite predictive of near-term outcomes, but become useless for longer horizon forecasts. In my experience, when information like this is available, no time-series technique that is unable to leverage this information would beat even simple regressions for short term horizon forecasts.
- lukas_b 3y agoThis looks very interesting! I'm trying to understand if the flattening technique might work for my ts. It's structured as follows: At each time step t, I have an m by n data matrix. The value for m (rows) varies per time step. n stays constant and represents the features. And i want to predict one of the n values. (In this case, t represents a single day, m (rows) represent the people that entered a store on that day, and n (cols) represent various features of the people. I want to predict one of those features, given the others.) The fact that it's a time series matters, because i expect the relationship to change over time. For instance some feature n[x] (person wears a yellow shirt) might be correlated with my target feature n[y] (person steals) but only in the summer. would it be possible to flatten this too? What would that look like?
- hackerlight 3y agoThey flatten the time and variate dimensions into a single 1D vector. So it can handle arbitrary numbers of features.
- longdog 3y agoInteresting, but I'm very skeptical. There are over a dozen transformers-based foundation time series model released in the past year and without fail, every one of them claims to be at or near SOTA. For example: - Time-LLM (https://arxiv.org/abs/2310.01728 https://arxiv.org/abs/2310.01728) - Lag-Llama (https://arxiv.org/abs/2310.08278 https://arxiv.org/abs/2310.08278) - UniTime (https://arxiv.org/abs/2310.09751 https://arxiv.org/abs/2310.09751) - TEMPO (https://arxiv.org/abs/2310.04948 https://arxiv.org/abs/2310.04948) - TimeGPT (https://arxiv.org/abs/2310.03589 https://arxiv.org/abs/2310.03589) - TimesFM (https://arxiv.org/html/2310.10688v2 https://arxiv.org/html/2310.10688v2) - GPT4TS (https://arxiv.org/pdf/2308.08469.pdf https://arxiv.org/pdf/2308.08469.pdf) Yet not a SINGLE transformer-based model I've managed to successfully run has beaten gradient boosted tree models on my use case (economic forecasting). To be honest I believe these foundational models are all vastly overfit. There's basically only 2 benchmarking sets that are ever used in time series (the Monash set and the M-competition set), so it'd be easy to overtune a model just to perform well on these. I would love to see someone make a broader set of varied benchmarks and have an independent third party do these evaluations like with LLM leaderboards. Otherwise I assume all published benchmarks are 100% meaningless and gamed.
- dcl 3y agoWhy would you expect anything to work well for economic forecasting :p
- donbreo 3y agoJamie pull up the article that proves none of the published models work well with economic forecasting
- rokkitmensch 3y agoI'm so sad. This hilarious comment is languishing in the doldrums.
- idiotsecant 3y ago
- wenc 3y agoThey should sign up for the next Makridakis forecasting competition. https://en.wikipedia.org/wiki/Makridakis_Competitions https://en.wikipedia.org/wiki/Makridakis_Competitions Makridakis and Hibon reached the sad conclusion that "statistically sophisticated and complex methods do not necessarily provide more accurate forecasts than simpler ones."
- hcarlens 3y agoThat was true in the first Makridakis competition ("M1") in 1982, and possibly until M4 in 2018, but both M5 and M6 were won by what would generally be considered relatively sophisticated methods (e.g. LightGBM). The Wikipedia article doesn't have that much detail on M5 or M6, but the M5 papers are in the International Journal of Forecasting[1] and M6 should be published later this year (there's already a preprint on arxiv [2]). I recently spent some time looking into the history and results of the M competitions and had a chance to speak to Professor Makridakis about them, as well as the winners of each of the M6 competition tracks [3]. While the methods have become more sophisticated, some conclusions from M1 still seem to hold: in particular, that there is no overall "best" method, and that the winning method tends to be different for different types of data, time horizons, and evaluation metrics. [1]: https://www.sciencedirect.com/science/article/pii/S0169207021001874?via%3Dihub https://www.sciencedirect.com/science/article/pii/S016920702... [2]: https://arxiv.org/abs/2310.13357 https://arxiv.org/abs/2310.13357 [3]: https://mlcontests.com/state-of-competitive-machine-learning-2023/#m6-financial-forecasting https://mlcontests.com/state-of-competitive-machine-learning...
- vermorel 3y agoOur basic low-dimensional parametric model landed No1 at the SKU level at the M5, see my lecture https://www.lokad.com/tv/2022/1/5/no1-at-the-sku-level-in-the-m5-forecasting-competition/ https://www.lokad.com/tv/2022/1/5/no1-at-the-sku-level-in-th... (more references at the bottom)
- hcarlens 3y agoInteresting, thanks for sharing!
- thelastbender12 3y agoI'm curious where universal forecasting models are most useful. It is technically fascinating but forecasting specifically seems like a domain where you'd want interpretable modeling - you use it for big-value problems and it significantly affects your action/policy. So, the tradeoff between performance and model simplicity should lean towards the latter?
- jdowner 3y agoSo I am not alone! There seem so few people who hold this view these days.
- paul80808 3y agoSame for my shop - we manage a large pool of cost driven by partially forcastable factors; we've repeatedly rejected methods purely on explainability grounds. Our accountability requirements do not allow us to point the finger at an LLM if we get it wrong.
- melondonkey 3y agoI know. Here I am modeling my data generating process like a chump.
- xotesos 3y ago[dead]
- laylower 3y agoShow us how it performs against other models on the M3, M4 and M5 competition. This is the gold standard of forecasting tools. Moirai stands for fates [https://en.wikipedia.org/wiki/Moirai https://en.wikipedia.org/wiki/Moirai] in Greek mythology
- tsurba 3y agoVery cool that the dataset and model weights are open right away! This paper also doesn't have a bunch of weird architectural choices pulled out of nowhere like the other TS foundation models recently. Looks like it will actually be useful, thank you! Maybe I will actually get to do representation learning for TS during my PhD. As a sidenote/rant, it would be nice if all supervised TS benchmarks included "DLinear + RevIN" as the standard baseline, as in my experiments it tends to get about the same performance as all other new SOTA forecasting models. Most papers compare to the linear model without RevIN while they themselves use it, and only beat it because of that :) And in any case supervised training of transformers from scratch on datasets having less than 1M points is just stupid (so less raw data than a single image?). Less than 1B is still at least mildly stupid. Here of course the angle is zero-shot so its somewhat excused from this, but it still would be interesting whether it can beat that supervised model combination.
- pbronez 3y agoReferences this paper on Time Series transformers. First I’ve seen someone apply transformers to time series specifically. Very curious how well this might work for low-frequency events. https://arxiv.org/abs/2402.02592 https://arxiv.org/abs/2402.02592
- melondonkey 3y agoOne detail I don’t really understand is the low-variance normal component of the target mixture. Would be curious to see from the weights how often that was used
- magundu 3y agoAny one tried this for Prometheus metrics?