23 ms·
"The most frustrating paper: I have true hate for the authors of this paper: "A deep learning framework for financial time series using stacked autoencoders an
by linux_devil 7y ago
"The most frustrating paper:
I have true hate for the authors of this paper: "A deep learning framework for financial time series using stacked autoencoders and long-short term memory". Probably the most complex AND vague in terms of methodology and after weeks trying to reproduce their results (and failing) I figured out that they were leaking future data into their training set (this also happens more than you'd think)."
- Not sure how author tried to implement it , but is this not how you train LSTM networks by feeding t+1 data back into the cell again to predict t+2 data. It will be easier if author made it open source as well
- laichzeit0 7y agoLeaking future data in would be using t+1 for t, e.g. something like a bi-directional LSTM. I assume he means the actual training dataset had some kind of signal in the data that was also in the test data.
- sgt101 7y agoPeople do this by doing things like testing that their features contain information in both the training and test set. Because they are not exposing the data directly to the classifier they think that they haven't compromised the test set - but what they have done is increased the chances of a chance correlation.
- jensgrud 7y agoI wrote my thesis last year comparing different RNNs against each other using this exact paper as baseline and basically concluded that you would be better off predicting the price yesterday than using their results. Authors did not respond when prompted for implementation details or comments. Overall, concluded that amongst RNNs the GRU architecture proved most favorable but still would not outperform simple stochastic models of the financial industry toolbox. You can check it out here https://github.com/jensgrud/financial-forecasting-lstm/tree/master/report https://github.com/jensgrud/financial-forecasting-lstm/tree/...