7 ms·
> In recent years, transformer-based models have gained prominence in multivariate long-term time series forecasting Prominence, yes. But are they generally be
by carbocation 2y ago
> In recent years, transformer-based models have gained prominence in multivariate long-term time series forecasting
Prominence, yes. But are they generally better than non-deep learning models? My understanding was that this is not the case, but I don't follow this field closely.
- Pandabob 2y agoWhile I don't have firsthand experience with these models, I recently discussed this topic with a friend who has used tree-based models like XGBoost for time series analysis. They noted that transformer-based architectures tend to yield decent performance on time series tasks with relatively little effort compared to tree models. From what I understood, tree-based models can usually outperform transformers when given sufficient parameter tuning. However, models like TimeGPT offer decent performance without extensive tuning, making them an attractive option for quicker implementations.
- techwizrd 2y agoIn my aviation safety work, deep learning outperforms traditional non-DL models for multivariate time-series forecasting. Between deep learning models, I've had a wide variance in performance between transformers, Bi-LSTMs, regular MLPs, VAEs, and so on.
- theLiminator 2y agoWhat's your go-to model that generally performs well with little tuning?
- techwizrd 2y agoIf you have short time-series with low variance, noise and outliers, strong prior knowledge, or limited resources to train and maintain a model, I would stick with simpler traditional models. If DL is a good fit for your use-case, then I tend to like transformers or combining CNNs with recurrent models (e.g., BiGRU, GRU, BiLSTM, LSTM) and optional attention.
- montereynack 2y agoSeconding the other question, would be curious to know
- ramon156 2y agoNow take into account that it has to be lightweight and DL falls shirt
- nerdponx 2y agoWhat are you doing in aviation safety that requires time series modeling? Weather?
- all2 2y agoMy best guess would be accident occurrence prediction.
- deleted 2y ago[deleted]
- dongobread 2y agoFrom experience in payments/spending forecasting, I've found that deep learning generally underperform gradient-boosted tree models. Deep learning models tend to be good at learning seasonality but do not handle complex trends or shocks very well. Economic/financial data tends to have straightforward seasonality with complex trends, so deep learning tends to do quite poorly. I do agree with this paper - all of the good deep learning time series architectures I've tried are simple extensions of MLPs or RNNs (e.g. DeepAR or N-BEATS). The transformer-based architectures I've used have been absolutely awful, especially the endless stream of transformer-based "foundational models" that are coming out these days.
- sigmoid10 2y agoTransformers are just MLPs with extra steps. So in theory they should be just as powerful. The problem with transformers is simultaneously their big advantage: They scale extremely well with larger networks and more training data. Better so than any other architecture out there. So if you had enormous datasets and unlimited compute budget, you could probably do amazing things in this regard as well. But if you're just a mortal data scientist without extra funding, you will be better off with more traditional approaches.
- dongobread 2y agoI think what you say is true when comparing transformers to CNNs/RNNs, but not to MLPs. Transformers, RNNs, and CNNs are all techniques to reduce parameter count compared to a pure-MLP model. If you took a transformer model and replaced each self-attention layer with a linear layer+activation function, you'd have a pure MLP model that can model every relationship the transformer does, but can model more possible relationships as well (but at the cost of tons more parameters). MLPs are more powerful/scalable but transformers are more efficient. Compared to MLPs, transformers save on parameter count by skimping on the number of parameters devoted to modeling the relationship between tokens. This works in language modeling, where relationships between tokens isn't that important - you can jumble up the words in this sentence and it still mostly makes sense. This doesn't work in time series, where relationships between tokens (timesteps) is the most important thing of all. The LTSF paper linked in the OP paper also mentions this same problem: https://arxiv.org/pdf/2205.13504 https://arxiv.org/pdf/2205.13504 (see section 1)
- rjurney 2y agoThey aren’t so hot, but recent efforts at transfer learning were promising.
- svnt 2y agoThe paper says this in the next paragraph. xLSTMTime is not transformer-based either.
- deleted 2y ago[deleted]