5 ms·
I do not have a horse in the race, but it is interesting to see open source comparisons to traditional timeseries strategies: https://github.com/Nixtla/nixtla/t
by izyda 3y ago
I do not have a horse in the race, but it is interesting to see open source comparisons to traditional timeseries strategies: https://github.com/Nixtla/nixtla/tree/main/experiments/amazon-chronos https://github.com/Nixtla/nixtla/tree/main/experiments/amazo...
In general, the M-Competitions (https://forecasters.org/resources/time-series-data/ https://forecasters.org/resources/time-series-data/), the olympics of timeseries forecasting, have proven frustrating for ML methods... linear models do shockingly well and the ML models that have won, generally seem to be variants of older tree-based methods (ie. LightGBM is a favorite).
Will be interesting to see whether the Transformer architecture ends up making real progress here.
- one_buggy_boi 3y agoAre these models high risk because of their lack of interpratability? Specialized models like temporal fusion transformers attempt to solve this but in practice I'm seeing folks torn apart when defending transformers against model risk committees within organizations that are mature enough to have them.
- tomrod 3y agoInterpretability is just one pillar to satisfy in AI governance. You have build submodels to assist with interpreting black box main prediction models.
- rdedev 3y agoIs there a way to directly train transformer models to output embeddings that could help tree based models downstream? For tabular data tree based models seems to be the best but I feel like foundational models could help them in some way
- wenc 3y agoThey are comparing a non-ensembled transformer model with an ensemble of simple linear models. It's not surprising that the ensemble models of linear time series models will do well, since ensembles optimize for the bias-variance trade-off. Transformer/ML models by themselves have a tendency to overfit past patterns. They pick up more signal in the patterns, but they also pick up spurious patterns. They're low bias but high variance. It would be more interesting to compare an ensemble of transformer models with an ensemble of linear models to see which is more accurate. (that said, it's pretty impressive that an ensemble of simple linear models can beat a large scale transformer model -- this tells me the domain being forecast has a high degree of variance, which transformer models by themselves don't do well on.)
- gradascent 3y agofyi I think you have bias and variance the wrong way around. Over-fitting indicates high variance
- wenc 3y agoThank you for catching that. Corrected.
- hackerlight 3y ago> ensemble of transformer models Isn't that just dropout?
- mikkom 3y agoNo. Why do you think so?
- hackerlight 3y agoGeoffrey Hinton describes dropout that way. It's like you're training different nets each time dropout changes.
- wenc 3y agoDropout is different from ensembles. It is a regularization method. It might look like an ensemble because you’re selecting different subsets but ensembles combine different independent models rather than just subset models.
- wenc 3y agoThat said random forests are an internal ensemble, so I guess that could work. In my mind an ensemble is like a committee. For it to be effective, each member should be independent (able to pick up different signals) and have a greater than random chance of being correct.
- 3y ago