16 ms·
MIT computer scientists can predict the price of Bitcoin
- jacobwcarlson 12y agoThe paper states that the strategy was simulated with live data and makes no mention of slippage. I've never traded bitcoin so I'm not sure how difficult it is to get fills, but that along with spreads are non-trivial components of real trading.
- dnautics 12y agoit would have been better if they set up an wallet and put bitcoin in it and demonstrated the trades that were automatically executed by this wallet. Since bitcoin is a ledger history, we could be 100% certain that SOME algorithm 'did the right thing', although we couldn't be certain that they spun up more than one algorithm and then just showed us the best one.
- ne0n 12y ago...That's not how Bitcoin trading works. In order to trade, your bitcoins need to be deposited into an exchange. Trades are not recorded on bitcoin's public ledger because there isn't a transaction for every trade. Bitcoin cannot currently handle that many transactions.
- gnaritas 12y agoTrading doesn't happen on the blockchain, it's far too slow.
- api 12y ago... for the next 15 minutes. When you predict the future of a market, you change the future of that market. People start investing on the basis of your predictions and whatever opportunity for profit you found is closed. This is why HFT people iterate constantly and also why they put their servers as physically close to the market as possible.
- mattfrommars 12y ago> This is why HFT people iterate constantly and also why they put their servers as physically close to the market as possible. Do you happen to have read this :- http://www.nytimes.com/2014/04/06/magazine/flash-boys-michael-lewis.html?_r=4 http://www.nytimes.com/2014/04/06/magazine/flash-boys-michae...
- barisser 12y agoThis seems like massive historical overfit, which can lead to arbitrarily precise fit, but no predictive capability. Any model, if given enough parameters, can be made to match historical data to an arbitrary degree. I also run several Bitcoin bots. I can tell you that slippage is not insignificant. If you make transactions every ~10 seconds and incur 0.1% fees each time, this is an extremely significant effect in aggregate. Also bid-ask spreads, while usually small, often aren't in periods of high volume.
- api 12y ago"Any model, if given enough parameters, can be made to match historical data to an arbitrary degree." I've known this forever but for some reason haven't heard this precise statement of it. Thanks. Reductio ad absurdium: imagine a model where the number of parameters equals the number of data points. Obviously that model will have perfect fit. Predicting the future is hard. Predicting the future without a causal understanding of the system is epistemologically questionable.
- arno_v 12y agoThis picture shows quite nicely what might happen when having too many parameters (or too little data): http://machinelearningac.files.wordpress.com/2011/10/polynomials.png http://machinelearningac.files.wordpress.com/2011/10/polynom...
- vmarsy 12y agoIn red is your model whereas in green is the real one, M being the number of parameters. The technical term for the last one is "overfitting" if I remember correctly. But in the case you have an enormous amount of data, it is unlikely to happen. It reminds me of this awesome course: https://www.coursera.org/course/ml https://www.coursera.org/course/ml edit: The parent's parent's parent mention overfit for the MIT work, I don't think it'd be the case if you have that amount of data in hands
- 12y ago
- joshdance 12y agoNo they can't. If they could they wouldn't tell anyone, and they would make millions (billions?) of dollars.
- Symmetry 12y agoThey should have made more money rather than publishing more quickly. It used to be possible to do these sorts of things to the stock market but when these sorts of regularities are discovered the process of exploiting them also eliminates them once enough money is being made. Heck, a major trading firm got started by noticing that stocks went down on the weekend (and of course they don't any more).
- HockeyPlayer 12y agoThey didn't include any discussion of: 1) execution (are they expecting to buy on the bid and sell the offer?). 2) commissions. They only made 3,362 yuan on 2,872 trades. A yuan is about 12 cents, so they are making 15 cents USD per trade. A .1% commission would cost them roughly 5 yuan per trade, but they are only making 1.17 yuan/trade.
- princeb 12y agoOKcoin (the exchange they relied on in the paper for data) claims 0% commissions and a bid-ask spread of just 0.04 RMB. how real is that, I have no idea. but if you relied on this information you could believe that you can still take a few additional spreads worth of slippage (in addition to crossing) to compensate for execution latency and still come out very far ahead
- rdmcfee 12y agoThis kind of innovation is cool, but it's a zero sum game. The bitcoin markets are already driven by competing bots. Their profits will be reduced as other bots iterate on their algorithms.
- joosters 12y agoPleas don't spout out 'zero sum' like you think it means something insightful here. Any marketplace is 'zero sum' if you think about it. There's a buyer and seller. So what?
- alexchamberlain 12y agoYou're assuming the exchanges don't take a cut...
- joosters 12y agooh, I'm not arguing that it is zero-sum, just that whether it is or not has absolutely no relevance here.
- pbhjpbhj 12y ago>Any marketplace is 'zero sum' if you think about it. vs >I'm not arguing that it is zero-sum So, what are you saying there then.
- Enzolangellotti 12y agoHow can the stock market be zero sum. What about dividends?
- tim333 12y agoZero sum markets are not really a thing. The normal phrase is 'zero sum game'. Buying and holding stocks is not a zero sum game - you make money. Trading stocks between speculators where the trading commissions exceed the dividends and long term market appreciation is a zero sum game. Same market, different games.
- Animats 12y agoThis is short-term trading. "Every two seconds they predicted the average price movement (on OKcoin) over the following 10 seconds. If the price movement was higher than a certain threshold, they bought a Bitcoin; if it was lower than the opposite threshold, they sold one; and if it was in-between, they did nothing." I don't see them allowing for commissions and fees. OKcoin, at peak, had a trading volume so high that it's generally considered to be fake - the exchange operators manipulating the price. What this group at MIT may have done is reverse-engineered the fake trade generation algorithm.
- rwissmann 12y agoPlus, from a quick glance, it looks like in-sample data with optimisation of parameters around scaling and holding periods. Fee-less trading on in-sample data. You can get the same return on futures markets with that.
- chollida1 12y ago> What this group at MIT may have done is reverse-engineered the fake trade generation algorithm. Just to be clear, there is nothing wrong with this. Infact, sitting around and reverse engineering what other traders are doing is what many funds do. I'm in this group so I"m happy to answer questions if anyone has any. > Every two seconds they predicted the average price movement (on OKcoin) over the following 10 seconds. If the price movement was higher than a certain threshold, they bought a Bitcoin; if it was lower than the opposite threshold, they sold one; and if it was in-between, they did nothing. To be clear, this is the core of what most HFT systems do these days. Consume many different factors, give each factor a decay factor to tell the system when the signal goes stale and distill them all into a value that says, but or sell or stay. It's worked well for Renaissance Technologies:) http://en.wikipedia.org/wiki/Renaissance_Technologies http://en.wikipedia.org/wiki/Renaissance_Technologies
- foobarqux 12y ago> Just to be clear, there is nothing wrong with this. Infact, sitting around and reverse engineering what other traders are doing is what many funds do. I'm in this group so I"m happy to answer questions if anyone has any. His point is that the posted trades were not in fact tradeable.
- sillysaurus3 12y agoI fixed the article's headline image: http://i.imgur.com/QVgcgNI.png http://i.imgur.com/QVgcgNI.png If everyone began using the paper's strategy, would the strategy still work? Also, the strategy seems less effective than portrayed in the news article. If you look at the "results" section, it seems like the profit flatlined shortly after starting, then had success due to some major trading event, then eventually flatlined again: http://i.imgur.com/CBjEjgo.png http://i.imgur.com/CBjEjgo.png Wouldn't it be more accurate to say "this strategy is effective under some very specific circumstances"? Also, does anyone know how equation 4 was derived? http://i.imgur.com/vkx8ZEC.png http://i.imgur.com/vkx8ZEC.png It seems like the key insight of the paper, but there's no mention of where it originated from. Is it a common equation in statistical modeling? I'd like to learn more about it. Does anyone have any suggested reading or coursework I should study?
- dash2 12y agoI think equation 4 is just derived from equation 3. The top bit counts every time before that y_i has taken a particular value y, and multiplies it by the distance of the x value then, x_i, from the value of x now. (Squared and exponentiated because this is the pdf of the normal distribution.) Here's an interpretation: if there were no noise, you might just count the number of times x took on a particular value and y took on a particular value, and divide that by the total number of times x took on that value. This would give you an empirical estimate of the prob of y given x. Because there's noise, they weight the counts by the pdf of the normal distribution of x - x_i. So, whenever x_i was close to current x, and y_i was a given y, that increases the probability of y occuring now.
- ISL 12y agoCan a HFT-knowledgeable commenter chime in on the viability of the Sharpe ratio here? From a physics perspective, it appears that the Sharpe ratio of 4.1 is roughly equivalent to a 4.1-sigma claim that their algorithm is better than random trading. I can't check easily, but I'd guess that the movement of Bitcoin prices isn't normally-distributed (looking at the paper's time series suggests that there's more low-frequency power there). If so, I'd guess that a more robust measure of the claim's significance would show it to be less significant. Put differently, I'd guess more than 1 in 15,000 random sets of 2872 trades (their number of trades) would yield comparable profit. Furthermore, a simple buy-and-hold would've yielded a 20+% return over the same period.
- lrm242 12y agoA sharpe of 4 is ok. Most "HFT" type strategies have sharpes so high they don't talk about sharpe anymore as it is meaningless. It is simply a different type of trading. See Virtu's pnl distribution in their S1. As a point of reference, Blair Hull is on record saying they are interested in nothing below a sharpe of 10 (r-finance talk, should be googleable).
- petercoolz 12y agoThanks for sharing. Link: http://www.nasdaq.com/markets/ipos/filing.ashx?filingid=9443224 http://www.nasdaq.com/markets/ipos/filing.ashx?filingid=9443...
- jff 12y agoCan they predict how to get real money back in exchange for their Bitcoins?
- xs 12y agoThose interested in automatic trading of bitcoin using algorithms should check out https://cryptotrader.org https://cryptotrader.org
- andrea_s 12y agoHow many of the people who are bashing the paper in this discussion are machine learning experts? Especially the overfitting crowd - looks like emotional attachment to one's favorite topics is not really impacted by said person's overall education and expertise.
- Houshalter 12y agoHere is the discussion on /r/MachineLearning: https://www.reddit.com/r/MachineLearning/comments/2jwpv1/mit_computer_scientists_can_predict_the_price_of/ https://www.reddit.com/r/MachineLearning/comments/2jwpv1/mit...
- xxcode 12y agoSLIPPAGE SLIPPAGE SLIPPAGE!!!
- minimax 12y agoThe problem with the paper is not overfit. They claim to have run their simulation with out of band ("live") data. The actual problem with the paper is that we have no idea if their simulator is any good, which means that their result (89% return in 50 days) could be totally bogus. In other words, we don't know if the actual bitcoin exchange would fill their orders at the same prices (if at all) as their simulator does. A decent simulator for a high-frequecy strategy like this is not trivial because you have to incorporate all the exchange behavior (documented or not) into your simulator, and then you have to validate your results by comparing your simulated results to the results of some actual trading. The fact that they spent none of the paper on the details of the simulator makes me extremely skeptical.
- latj 12y agoImagine if someone learned how to legally print money and their first action was to tell the world-- then I would be extremely skeptical.
- lrm242 12y agoPeople so often overlook the role that execution plays in trading. As time scales shrink the impact of execution increases, and indeed many HFT strategies are not profitable without top tier execution (both in terms of fee schedules and technology). This is also one of the places that most academics fail when analysing a trading strategy. They make typical "assume a frictionless surface" types of assumptions that break down in a real market. To your point, high-frequency simulation is damn hard and it is highly likely that they failed here and have tainted their results with bad simulations. Most of these papers would be well served to avoid dollars and cents and simply analyse the statistical qualities of their signal to their target over time. That is the first step in this business, anyway. Simulations only happen once the signal has been through the wringer.
- hft_throwaway 12y agoI see no mention of the spread, costs or even ability to put on a short position, exchange lag variability (huge issue when simulating even on modern exchanges, let alone fly-by-night bitcoin markets). Additionally, it is very easy to find trend-following signals that "work" but break down when conditions change very quickly. That seems to make up most of the prediction, along with an order book imbalance signal that might be more stable. I think a better approach would be looking at order book features and lags vs. other markets. The costs are high on BTC markets so it would be pretty tough to overcome those though. Anyone with experience to do this is probably doing it somewhere more lucrative. I think BTC markets only trade a few million USD a day.
- zoba 12y agoHe says "Give me your money and I’d be happy to invest it for you." Alright, I'll do it... just tell me how.
- crimsonalucard 12y agodoesn't predicting the price change the price? Just like how knowing the future changes the future.
- Houshalter 12y agoSometime in 2013, before the bitcoin prices exploded, I downloaded some bitcoin historical price data and ran symbolic regression on it with Eureqa. It came up with a formula that fit the observed data fairly well, and wasn't very complicated. But when I extrapolated it forward a few months, it predicted the price would explode to unreasonable levels. I was disappointed and threw it away, assuming that it must be wrong.
- runeks 12y agoSuccess in predicting markets is measured in profit. This research team should start a company that offers a service that allows users to deposit bitcoins, which the company then invests according to their alleged predictions, and then pay interest on deposits, and keep a part of the profit for themselves. Doubling the initial investment one time is one thing, but this hypothetical company being able to double its investment every 50 days for years is something else. I doubt they can. A doubling every 50 days is x160 every year. I think claims of being able to predict market prices should be met with great skepticism. Especially prices of easily traded commodities, including bitcoins. The only proper measure of an ability to predict market prices is profit, because profit also measures the extent of the predictions: how much can you move the market (by trading according to your predictions) until you can no longer predict what will happen? Obviously, there's a limit. No one can extract unlimited profit from any market. So there definitely is a limit to how much you can earn from your algorithm. If you can earn 10% p.a. on an investment of maximum $5000, your algorithm isn't really worth much. If you can earn 1000% p.a. on an investment of up to $100M, your algorithm is great. But without knowing these figures we really only have a claim, seemingly a claim of them being able to make a lot of money, but choosing not to do so.
- deleted 12y ago[deleted]
- stokk 12y agoEver thought how our big data overlords (Google, Facebook, MS, Twitter, etc) just need to check for correlations between their users' data input and stock exchange movement? They own the stock "matrix".
- liaboc 12y agoYou mean something like this? http://financeai.com/stock/nyse/twtr http://financeai.com/stock/nyse/twtr
- liaboc 12y agoThe fact is anyone can *Predict the price of bitcoin, but will you actually put your own money to trade? Like what we did here, http://financeai.com/forex/btc http://financeai.com/forex/btc it does show some degree of correlation between sentiment and the price.
- Belmont1 12y ago89% over 2 months. Not bad. My algorithm did 8,000% over 6 months using real money on real exchanges and I have the trade history to back it up.
- martin1975 12y agoJordan Belfort, is that you?
- damian2000 12y agoWhat exchanges do you recommend for bot trading? ... I've dabbled a bit trading manually on bitfinex but that's all.
- jaekwon 12y agoBy publishing the paper you essentially invalidate it, as people take advantage of it. Happens time and time again in any market.
- deleted 12y ago[deleted]
- Macuyiko 12y agoI'm late to comment, but something which I'd like to point out is that this is done by the same team behind the Twitter trending topic prediction technique from a few years back, as mentioned also in the article [1]. When their Twitter technique was released, I spend a few weeks reading through Nikolov's PhD thesis (the advisor gets most of the the fame in the press articles but Nikolov's thesis has all the details) and trying to implement it in R. My observations at the time: extremely simple algorithm which would be shot down by most peer reviewers for being not very novel (the affiliation helps a lot here). That said, I believe greatly in pragmatism, and the approach was actually working well. What I did find out however is that their was a great deal of data selection and pre-processing involved making the approach hard to implement in a real-life, real-time setup. I get similar feelings from this work. [1]: http://newsoffice.mit.edu/2012/predicting-twitter-trending-topics-1101 http://newsoffice.mit.edu/2012/predicting-twitter-trending-t...
- atomroflbomber 12y agoCould you specify what kind of data selection and pre-processing was required and why exactly it was hard to implement it in a real-life setup?