10 ms·
Time Series Prediction Using LSTM Deep Neural Networks
- lettergram 8y agoI've built quite a few of these kinds of models. The real trick is to compare it against other methods AND to properly split A LOT of data. In many cases, (depending on the input data) a random walk does roughly as well as "predicting". This is because signal data (such as stock data) often just follow a random (or seemingly random) trend.
- jarym 8y agoWhy does everyone naively try to predict price? No ‘traders’ are interested in predicting it - what traders do is identify good locations to enter or exit the market. I.e. places with defined risk where you will know if you’re wrong if it goes against you by x% while you expect a y% gain if you’re right AND y>x is worth more than the number of times you’re wrong. The types of Algos that work well for this are edge identification ones - I know this because I am (not as well as I’d like) successfully doing it. LSTMs haven’t performed so well for me in this task but non-NN algos have. CNNs however were promising but didn’t match what I’d come up with - still searching for the holy grail that’ll make me rich!
- rahimnathwani 8y agoCan you point to any good sources for this? A search for 'edge identification stocks' yielded mostly irrelevant results.
- jarym 8y agoFor example this blog post (not mine): http://jon.io/machine-beats-human-using-machine-learning-in-forex.html http://jon.io/machine-beats-human-using-machine-learning-in-... Talks about the Mean Shift algorithm described here: https://en.wikipedia.org/wiki/Mean_shift https://en.wikipedia.org/wiki/Mean_shift
- polskibus 8y agoCan you share some pointers to more reading material on algos you consider more successful?
- foxes 8y agoBetter, why do people think they can predict long term dynamics of what is basically a chaotic system? I think finding some small local dynamics is fine, but applying a neural network to try say something about global long term dynamics complete garbage - ie long term weather simulation, stock markets etc. Chaotic systems can be deterministic, just that you will never be able to accurately measure all the variables to make a long term prediction accurately. In the weather example, people know the equations that approximate how it works. Why value does a neural network bring? Knowing the equations is better understanding.
- jkabrg 8y agoThere was a Quanta Magazine article talking about predicting the evolution of a flame-front using ML. The ML algorithm remained accurate for eight Lyapunov intervals; eight times longer than the previous SOTA. https://www.quantamagazine.org/machine-learnings-amazing-ability-to-predict-chaos-20180418/ https://www.quantamagazine.org/machine-learnings-amazing-abi...
- foxes 8y agoThe only real argument I think you can make is that it might be more efficient to have a neural network quickly spit out an approximate solution instead of solving the actual equations. But if you have the time having the actual equations is more valuable?
- PeterisP 8y agoAssuming that "having the actual equations" is possible, which in market predictions it most likely is not.
- bougiefever 8y agoIs this what you are talking about? https://www.dukascopy.com/fxcomm/fx-article-contest/?Automated-High-Frequency-Trading-With=&action=read&id=1835&language=en https://www.dukascopy.com/fxcomm/fx-article-contest/?Automat...
- cmroanirgo 8y ago> Why does everyone naively try to predict price? No ‘traders’ are interested in predicting it - what traders do is identify good locations to enter or exit the market. I agree. I've built many systems in this area, but it wasn't until I started working in the Indian market (>10 yrs ago) that it became abundantly clear that trying to calculate the long/shorts signals using historical (/time series) data was a waste of time. (And yet my primary role was to provide tools that did exactly that). Back then, in the indian market, you could see that most of the stocks, although skyrocketing upwards, all followed the slow vs fast moving averages to buy and sell! Back then, they weren't looking at RSI, stochastics, support lines, etc, etc. It was crazily predictable...but over time it was really interesting to see it become more haphazard and like western stocks. That is, the fundamentals came into play and as you say, the traders began to use other metrics to buy and sell.
- person_of_color 8y agoHow much did you make?
- leeuwnhawk 8y ago> I've built many systems in this area, but it wasn't until I started working in the Indian market (>10 yrs ago) that it became abundantly clear that trying to calculate the long/shorts signals using historical (/time series) data was a waste of time. (And yet my primary role was to provide tools that did exactly that). I'm currently working on building similar tools in my area of work for the Indian market and would really appreciate if you could shed some more light into the things you learned from your experience in working in this domain.
- beagle3 8y ago.. because you buy at the price, and sell at the price (spread and fees ignored for now). Which means, regardless of your philosophy, you are predicting a price change - a long signal is a prediction for positive price change; a short signal is a prediction for a negative price change. If that wasn’t true, your system would not be able to profit. Predicting price change and predicting price are semantically equivalent, although a specific algorithm might be better at one than the other.
- jarym 8y agoPredicting price means you’re predicting one variable with no idea of hot likely you are to be wrong and how wrong you’re likely to be and says nothing of where your expectations are for price to go after. It is semantically different to say: if price goes to Y then you have odds that it will then go to Target 1 and then slightly lower odds it goes to Target 2.
- beagle3 8y agoDo you long now, or short now, or neither? If you go long, you’ve predicted price goes higher. Regardless of what you think about the entire path. Mathematically nothing else makes sense.
- qeternity 8y agoGiven less than 100% certainty, traders don't want to predict price, they want to predict future distribution of price over some time period. Source: hedge fund trader
- beagle3 8y agoTrue. That’s still considered a prediction of price among my trading colleagues.
- noelsusman 8y agoYou want to predict price, but a price prediction is useless without an accurate estimate of the error of your prediction.
- glial 8y agoSeems to me that this is almost dangerous unless the uncertainty (and therefore confidence) of the prediction can be quantified.
- RA_Fisher 8y agoYep, it is dangerous. If you're not quantifying uncertainty, you can't make safe predictions. I think this is reason for the obsession with "data cleaning" in the ML community, "outliers" aka rare observations sink general models.
- fooker 8y agoSo, curve fitting?
- f00_ 8y agoexactly, Judea Pearl's The Book of Why opened my eyes to the fact that most of what happens in machine learning is really just curve fitting It connected with what i've heard Chomsky say about trying to develop laws of physics by filming what's happening outside the window. We need to do experiments and interventions to learn the dynamics of a system "What do you think the role is, if any, of other uses of so-called big data? [...] NOAM CHOMSKY: It’s more complicated than that. Let’s go back to the early days of modern physics: Galileo, Newton, and so on. They did not organize data. If they had, they could never have reached the laws of nature. You couldn’t establish the law of falling bodies, what we all learn in high school, by simply accumulating data from videotapes of what’s happening outside the window. What they did was study highly idealized situations, such as balls rolling down frictionless planes. Much of what they did were actually thought experiments. Now let’s go to linguistics. Among the interesting questions that we ask are, for example, what’s the nature of ECP violations? You can look at 10 billion articles from the Wall Street Journal, and you won’t find any examples of ECP violations. It’s an interesting theory-determined question that tells you something about the nature of language, just as rolling a ball down an inclined plane is something that tells you about the laws of nature. Scientists use data, of course. But theory-driven experimental investigation has been the nature of the sciences for the last 500 years. In linguistics we all know that the kind of phenomena that we inquire about are often exotic. They are phenomena that almost never occur. In fact, those are the most interesting phenomena, because they lead you directly to fundamental principles. You could look at data forever, and you’d never figure out the laws, the rules, that are structure dependent. Let alone figure out why. And somehow that’s missed by the Silicon Valley approach of just studying masses of data and hoping something will come out. It doesn’t work in the sciences, and it doesn’t work here." - https://www.rochester.edu/newscenter/conversations-on-linguistics-and-politics-with-noam-chomsky-152592/ https://www.rochester.edu/newscenter/conversations-on-lingui... It is actually a really interesting subject, marketing people doing a/b tests for ads/features seem at least a little closer to the experimental ideal, not just fitting curves to data For further reading, I'd recommend the epilogue of Casuality (Pearl 2000), it's from a 1996 lecture at UCLA: - http://bayes.cs.ucla.edu/BOOK-2K/causality2-epilogue.pdf http://bayes.cs.ucla.edu/BOOK-2K/causality2-epilogue.pdf
- shawn 8y agoIf anyone is looking to get into machine learning, I've found "Introduction to Data Mining" very useful: https://news.ycombinator.com/item?id=17808349 https://news.ycombinator.com/item?id=17808349 First edition: http://www.uokufa.edu.iq/staff/ehsanali/Tan.pdf http://www.uokufa.edu.iq/staff/ehsanali/Tan.pdf Also see "mining of massive datasets" usually available at this link, but it seems to be down: http://infolab.stanford.edu/~ullman/mmds/book.pdf http://infolab.stanford.edu/~ullman/mmds/book.pdf Which leads me to another point: Many of these books cost $100+. If you don't have those kind of resources, try Library Genesis. It's been very helpful for getting started.
- daviddumenil 8y agoCould this approach be applied to a metric monitoring framework to give earlier/more accurate notifications if when a threshold would be crossed? Typically these are triggered when e.g. 90% of a threshold has been crossed.
- cshenton 8y agoFor anyone considering this, LSTM only starts to pay off if you have many many time series. For a single time series like this one you’re better off using classical time series approaches like ARIMA or other Gaussian state space models.
- djhworld 8y agoI'm currently learning machine learning at the most basic level, this is the sort of stuff I want to work towards though I deal with time series data a lot at work, I work in broadcasting/media and 99% of the time the data is fairly "predictable" and follows a regular daily pattern, peppered with the odd spikes during big, unpredicatble news events.
- dafrie 8y agoA year ago, the original blog post [1] (it was just recently updated, which is now the one linked here on HN) helped me on a semester thesis, where I quite successfully used LSTM for short-term electricity load forecasting, which also has very strong daily, weekly and seasonal patterns. I used multiple features/variables such as calendar and weather data and found the LSTM models to easily beat ARIMA/TBATS forecasts. You can find the code repo on my Github link [2], but please bear with the code quality. I only have an economics background, so my coding experience is fairly limited :) [1] http://www.jakob-aungiers.com/articles/a/LSTM-Neural-Network-for-Time-Series-Prediction http://www.jakob-aungiers.com/articles/a/LSTM-Neural-Network... [2] https://github.com/dafrie/lstm-load-forecasting https://github.com/dafrie/lstm-load-forecasting
- md2be 8y agoTime series analysis requires the data to be stationary.
- dafrie 8y agoWell, I don't want to be pedantic, but don't you rather mean "Most TSA MODELS require data to be stationary"? My experience has been, that often practical TSA actually involves how to deal (testing, differencing, smoothing...) with non-stationarity, which is often not a trivial task...
- zwaps 8y agoI find it interesting that Computer Scientists are basically rediscovering statistics. Now when predicting time series, an issue is that most model (like ARIMA, GARCH etc.) are short-memory processes. When you look at the full-series prediction of LSTMs, you observe the same thing. So in terms of Time Series, Machine Learning is currently in the mid to late 80's compared to Financial Econometrics. So if you are a CS, you should now probably take a look at fractional GARCH models and incorporate this into the LSTM logic. If the statistic issues are the same, then this may give you that hot new paper.
- jkabrg 8y agoNassim Taleb had some negative things to say about GARCH. "GARCH does not work out of sample. It is a good story, but I was unable to use it in predicting squared deviations or mean deviations" I haven't found it in Rob J Hyndman's forecasting tutorial either. How does it fare in the Makridakis competitions?
- VHRanger 8y agoYou shouldn't listen to N. Taleb on technical matters. He's been a classic mold crank for the last decade or so when it comes to anything serious, relegated instead to writing fluffy books on whatever he thinks is important.
- zwaps 8y agoGARCH, like I said, is a short memory process and is inherently inadequate for (longer) out of sample predictions. Doing this is possible, but not really correct. Taleb is basically right, of course what he says is probably inflammatory and half wrong, as usual. Don't forget that most econometrics models are also concerned with identification and causality, less with prediction.
- RA_Fisher 8y agoIt's been amazing to watch CS (really the Python community, save statsmodels and patsy) discover statistics. For a while I thought perhaps it was me and statistics that was "behind." Over time I realized that it was mostly re-invention of old ideas: one-hot encoding = dummy variables, neural networks approximating polynomial regression, etc. I decided to double-down on statistics and it's really paid off. NN / random forests and the stats-founded but CS-led approaches are very general models. That leaves statisticians a big opening because a more specific model can be chosen to obtain more accurate predictions. These days I'm positioning myself to clean-up the messes / save broken ML models. Turns out [stats] theory is very practical. :-)
- GChevalier 8y agoI see here that original poster (OP) of the post tried to use many-to-one LSTMs instead of many-to-many LSTMs. I tell that first by looking at the charts. Then I saw the method named "predict_point_by_point" with the comment "Predict each timestep given the last sequence of true data, in effect only predicting 1 step ahead each time" in his code here: https://github.com/jaungiers/LSTM-Neural-Network-for-Time-Series-Prediction/blob/6aa5c5124ebfa405bf38bb6674871ab59d458b5c/core/model.py#L89 https://github.com/jaungiers/LSTM-Neural-Network-for-Time-Se... I strongly think the system would be better to perform many predictions at once instead, using seq2seq neural networks. The problem is properly explained here at the beginning of this other post: https://github.com/LukeTonin/keras-seq-2-seq-signal-prediction https://github.com/LukeTonin/keras-seq-2-seq-signal-predicti... This other post is, in turn, derived from my original project here doing seq2seq predictions with TensorFlow: https://github.com/guillaume-chevalier/seq2seq-signal-prediction https://github.com/guillaume-chevalier/seq2seq-signal-predic... OP also forgot to cite the image I made: https://en.wikipedia.org/wiki/Long_short-term_memory#/media/File:The_LSTM_cell.png https://en.wikipedia.org/wiki/Long_short-term_memory#/media/... Well, glad to see that some similar work as mine can get this much traction on HN. I would have loved to get this much traction when I did my post, too. Anyway, I would suggest OP to take a look at seq2seq, as it objectively performs better (and without the "laggy drift" visual effect observed as in OP's figure named "S&P500 multi-sequence prediction"). In other words, using many-to-one neural architectures creates some kind of feedback which doesn't happen with seq2seq which doesn't build on its own accumulated error. It has a decoder with different weights than the encoder, and can be deep (stacked).
- luanton 8y agohttps://news.ycombinator.com/item?id=17902967 https://news.ycombinator.com/item?id=17902967 The aim of this post is to explain why sequence to sequence models appear to perform better than "many to one" RNNs on signal prediction problems. It also describes an implementation of a sequence 2 sequence model using the Keras API.