4 ms·
This seems like massive historical overfit, which can lead to arbitrarily precise fit, but no predictive capability. Any model, if given enough parameters, can
by barisser 12y ago
This seems like massive historical overfit, which can lead to arbitrarily precise fit, but no predictive capability.
Any model, if given enough parameters, can be made to match historical data to an arbitrary degree.
I also run several Bitcoin bots. I can tell you that slippage is not insignificant. If you make transactions every ~10 seconds and incur 0.1% fees each time, this is an extremely significant effect in aggregate. Also bid-ask spreads, while usually small, often aren't in periods of high volume.
- api 12y ago"Any model, if given enough parameters, can be made to match historical data to an arbitrary degree." I've known this forever but for some reason haven't heard this precise statement of it. Thanks. Reductio ad absurdium: imagine a model where the number of parameters equals the number of data points. Obviously that model will have perfect fit. Predicting the future is hard. Predicting the future without a causal understanding of the system is epistemologically questionable.
- arno_v 12y agoThis picture shows quite nicely what might happen when having too many parameters (or too little data): http://machinelearningac.files.wordpress.com/2011/10/polynomials.png http://machinelearningac.files.wordpress.com/2011/10/polynom...
- vmarsy 12y agoIn red is your model whereas in green is the real one, M being the number of parameters. The technical term for the last one is "overfitting" if I remember correctly. But in the case you have an enormous amount of data, it is unlikely to happen. It reminds me of this awesome course: https://www.coursera.org/course/ml https://www.coursera.org/course/ml edit: The parent's parent's parent mention overfit for the MIT work, I don't think it'd be the case if you have that amount of data in hands
- barisser 12y agoHere's another way to think of it. If the parameter space for my model includes, let's say 10 binary decisions (which is very conservative), that's 1024 possible states of my model. If I tested all 1024 states against historical data, it is likely that some of them might do very well (depending on the general architecture of the model of course). What if I then selected the successful minority and held them up as clever strategies? Their success would very likely have been arbitrary. By basically brute-forcing enough strategies, I will inevitably come across some that were historically successful. But these same historically successful strategies are unlikely to outperform another random strategy in the future. It's not impossible you'll find a nugget of wisdom hidden from everyone else, just much less likely than the more simple explanation I'm offering. So to your point, it's not just the size of the parameter space versus the data set that matters. Brute-forcing the former alone will likely produce a deceptive minority of winners.
- Enzolangellotti 12y agoThere is a fun chapter on this topic in Jordan Ellenberg's latest book "How not to be wrong". It's called the "Baltimore stockbroker fraud".
- Houshalter 12y agoIt's entirely possible to overfit with enormous amounts of data. As people are now creating models with enormous numbers of parameters.
- jessaustin 12y ago"With four parameters I can fit an elephant, and with five I can make him wiggle his trunk." -- John von Neumann The green "real" signal in your picture is amusing when juxtaposed with the red "zero-noise-assumption" signal in the last frame. TBH this accounts for most of my distrust of e.g. climate modeling.
- dratman 12y ago"Predicting the future without a causal understanding of the system is epistemologically questionable." That is an intriguing assertion, but it is circular. One can only demonstrate a causal understanding of a system by making usefully accurate predictions about the system's behavior. To attempt that with a system consisting of market prices of tradeable securities is an exercise in frustration, because such markets do not operate by consistent, unchanging causal rules. In fact, financial markets are not systems at all in the usual sense, because their parts and connections are continually changing.
- foobarqux 12y agoCurious as to your returns, how they have changed over time and if you are doing this full time.
- leeber 12y agoAgreed. According to the paper: They trained the data, once, with 3-4 months of data from Feb to May. Then they tested the data, once, with ~6 weeks of data from May 6 - June 24. There was no cross validation involved in assessing the performance of this model. Using one subset for training, and one subset for testing and calling that conclusive is naive. No interesting conclusions can be taken away from this paper due to flawed methodology and failure to take into account real world variables like bid-ask spread, commissions, how quickly trades can be executed, whether orders will even be filled, etc.
- vasilipupkin 12y agoWell, they did trade it out of sample, so that probably means the model wasn't overfit. But, I agree, they probably used unrealistic execution assumptions
- beejiu 12y agoCross validation too uses training sets and test sets. This sort of time-ordered data will not be independent, so the prequential approach seems to be a more suitable approach to measuring forecast accuracy. (I haven't read the paper.)
- leeber 12y agoCross validation for time series, goes something like this: http://robjhyndman.com/hyndsight/crossvalidation/ http://robjhyndman.com/hyndsight/crossvalidation/ (near bottom of the article) Anyways, as somebody who has spent GIANT amounts of time experimenting with machine learning using market data (mainly stocks), I can give you TONS algorithms I've created that would show similar, even much higher, gains when you only test on 6 weeks of data. Been there done that. For example 6 weeks you make a 100% return, then in the next 3 weeks you suffer a 50% loss. Reality sets in...