10 ms·
Numerai – A hedge fund built by a global community of anonymous data scientists
- late2part 10y agoThis is darn cool
- zitterbewegung 10y agoWhy should I think that being anonymous should give some advantage ? I think it would be more of a disadvantage due to should I trust you? Also, how do I know these people are truly anonymous ?
- deleted 10y ago[deleted]
- daveguy 10y agoAnonymous in the sense that numerai does not ask for any personally identifying information. You could easily prove you are who you say you are in numerai. Also, I think the big benefit from a data scientist working on this is you can test methods with generic features, submit the results and get paid if they are good, but not submit any part of the methodology to any third party. If you can kill it on numerai then maybe you would consider buying data sources and apply your methods to your own data. Although you still don't know what the features are. It's the polar opposite of open source. The owners don't have to trust the data scientists. They evaluate their results against additional data.
- Houshalter 10y agoBut how do you trust the owners not to make fake accounts and only pay out to them?
- dharma1 10y agoThat's what I thought - if you are smashing it on numerai, wouldn't you be better off raising some capital for your own fund to scale it?
- jimfleming 10y agoSpent some time experimenting with Numerai. Really fun competition, clean (encrypted) dataset, and Bitcoin payouts. I wrote about my experience and open-sourced all of the models here[0] if you're looking to get started. [0] https://github.com/jimfleming/numerai https://github.com/jimfleming/numerai
- dharma1 10y agoAwesome. Thanks for the write up
- kowdermeister 10y ago"In December 2015, we created the world’s first encrypted data science tournament for stock market predictions. Since then, Numerai data scientists have submitted 13,350,675,598 equity price predictions. The most accurate and original machine learning models from the world’s best data scientists are synthesized into a collective artificial intelligence that controls the capital in Numerai’s hedge fund." What does this mean? What do they do?
- bigdubs 10y agoMonkeys throwing darts at the board. This brings to mind the Buffett hedge fund wager, where he invested in a vanguard s&p 500 tracking fund (VFIAX) and a hedge fund actively managed an equal amount, and Mr. Buffett ended up winning handily.
- throwaway98237 10y agoReally smart monkeys. And monkeys which will (hopefully) be self-correcting to converge on the bulls-eye. Most people don't realize that the "markets" are 49% random, 48% sentiment driven, and 3% fundamentals. If you approach the problem with that assumption held true, monkeys throwing darts isn't such a horrible mechanism for investing. See: "Monkeys Are Better Stockpickers Than You'd Think: Why dart-throwing primates demolish S&P 500 returns and most active fund managers don't even come close." http://www.barrons.com/articles/SB50001424053111903927604579634603846777722 http://www.barrons.com/articles/SB50001424053111903927604579...
- xg15 10y ago> Most people don't realize that the "markets" are 49% random [...] If you approach the problem with that assumption held true, monkeys throwing darts isn't such a horrible mechanism for investing. Maybe I haven't understood that part of data science but I never got why throwing more unpredictability on an already unpredictable data source would somehow make it more predictable.
- 10y ago
- bgitarts 10y agoThe training set just has anonymized features. Data scientists generally would like to know the nature of the data they are working with. Does the site at any point give access to the labeled featured?
- gravypod 10y agoCouldn't I make thousands of fake accounts and submit thousands of slightly different models. Then after a while I could push one or two of my stocks higher in my models artificially. If you represented a large enough % of the "data scientists" in this hedge fund you could make it look like your stock is "definitely" work investment. After they invest in your company, you could take off to the hills. Hell, you don't even need to be the owner of the company. This would be a great way to obtain large amounts of political sway/power. Like a company/want it to succeed for some agenda? Make it look better as an investment opportunity. Dislike a company? Well that stock is going to do horrible next quarter. It's also a self fulfilling prophesy. For all of you finance people out there, is what I am saying impossible or stupid? I hope I'm wrong otherwise this is a horrible idea.
- Hasz 10y agoIt sounds a bit like pump and dump, which is considered fraud. However, instead of being against hundreds or thousands of individual investors, you're going against one (obstinately) accredited investor, so you might be one firmer legal ground, as they should know better. It's probably against their TOS though.
- gravypod 10y agoAren't the "data scientists" anonymous? Login over tor, exert some basic opsec, and live in a country with no extradition treaty. Do it as a service for billionaires who want to play the never ending game of chess that is the the economy. For 100k per "tweak" you could make a lot of money and still give people rather significant influence over the world (especially if a LOT of people used this service). > pump and dump, which is considered fraud Yea it is definitely a pump and dump but if it makes you rich and you're in a country with no extradition treaty who cares? Remember the golden rule of the "elite class": laws are for poor people. As long as you get rich before anyone sees what you aredoing, and you "cant" be found, you're all good.
- jon_richards 10y agoIf they aren't managing correlation between different models there is an even bigger issue...
- exit 10y agoi suppose this could be, uh, useful for insider trading
- baccredited 10y agoIf I had a winning strategy why would I feed it to Numerai instead of instavest.com?
- yankoff 10y agoYou don't know if you have a winning strategy. You'd have to put your own money and take a risk to find out. Plus you'd have to take care of data and feature engineering yourself.
- numbas 10y agoBecause you don't have capital, risk-adversity, and access to the (expensive) dataset. You only have encrypted predictions, which are worthless for you.
- lordnacho 10y agoHedge fund guy here. - You need to know something about the domain in order to make sensible predictions. Is the data daily? Is it per second? Is it ticks? You can't build a sensible model if you don't know that, even if you have good predictions. Relative cost will vary a lot between timescales. - It matters what the features are. Maybe there's some clever reason why it doesn't, but until I hear why I'm going to take the ordinary view that some features are different in nature to others. For instance maybe one feature is volatility, a thing we typically model with GARCH, while another is some fundamental like P/E, which we'd incorporate some other way. - How are you executing the trades? It matters a lot whether you're click-trading through some broker API, automating via Excel, or running your own network of colo servers. Some things just aren't possible if you're too slow. - If you make the data encrypted, you'd better know very well what it represents. For instance, you might take all the closing prices of the LSE stocks on a given day as inputs. You can make analyses that are valid with that, and ones that aren't, because the data you've collected do not represent a snapshot of the market at a specific time. It might sound like it does, but it doesn't on deeper inspection (market opens and closes are not simultaneous). Does anyone know how it's going for them?
- dsacco 10y agoI don't know how it's going for Numerai, but I'd be interested in picking your brain a bit. If you don't want to put contact info in your profile, do you mind shooting me an email?
- numbas 10y agoYou don't really need to know what the features mean to make a good model. Even if only you knew the meaning of the features you'd still get beaten by someone using a blackbox approach. This is exactly what happened to the owner. He worked as a quant making predictions on this expensive dataset, and in the first tournament he was beaten by around 50 other modelers using anonymized features.
- fixxer 10y agoBased on the contest (weekly submission) and data (binary classification), it is my guess that they are running a long-short equity portfolio with weekly rebalancing. The features are known ahead of time and the dataset is not BIG, so I think some sort of fundamental approach. Im not optimistic...
- DrNuke 10y agoIn my limited experience in London at the trading level, they do not want collective intelligence at all: they do the bulk with algorithms at very low level latencies and want outliers (unpredictable singletons are less reproducible than the average from millions heads) to run high risk - high reward books.
- deleted 10y ago[deleted]
- delegate 10y ago"It's just a pure math problem. It's like a math competition. You don't need to know anything about finance, you don't know anything about hedge funds... you don't even have to speak English..." This is not a pure math problem. Eventually the outcome of all these models and predictions affects the stock prices and - if it becomes as successful as you hope - the economy as a whole. And the physical world: people, animals, plants, pollution, CO2 and so on. I would much rather see work in that area - than this data juggling deep learning bullshit which results in profits being paid out to a bunch of intelligent, greedy and unwise people. I just hope that when the intelligent people finally become wise, it won't be too late.
- werdna123 10y agoWe didn't start the fire It was always burning Since the world's been turning
- wyager 10y ago> than this data juggling deep learning bullshit which results in profits being paid out to a bunch of intelligent, greedy and unwise people. Holy anti-intellectualism, Batman! If you describe machine learning as "data juggling bullshit", this is very strong evidence that you simply don't understand what it is. This is an indictment of you, not of machine learning. Machine learning would more accurately be called "applied computational statistics" in 99% of cases. What makes you think that the people using applied statistics to make money are "unwise"? Based on your tone, I would guess that it's because what they're doing doesn't agree with your folk definition of "an honest day's work" or something like that. This isn't really a good criticism; it just means you don't see the utility of what they're doing, which requires some degree of abstract thinking about the market.
- pyrophane 10y agoI didn't read it as a criticism of machine learning or intellectuals, but rather of the many bright minds that are trying to make money off of the stock market by writing software to make better trades. This kind of activity doesn't create value or even, to my knowledge, liquidity. I believe he's suggesting that there are a lot of other areas where their knowledge and skills could be applied which would create some societal benefit as a byproduct.
- lasermike026 10y agoHow ominous. In today's news anonymous scientists form human genetics laboratory to improve the human species.
- DeBraid 10y agotldr - novel encryption method allows for sharing of data sets and better machine learning models. Aggregation of models into a portfolio offered as a hedge fund. Found this company interesting, read a bunch of blogs from them and tweetstormed: https://twitter.com/Royal_Arse/status/787725301908242432 https://twitter.com/Royal_Arse/status/787725301908242432
- cheiVia0 10y agoThings I do not understand: * How does a logloss relate to earnings? * They only receive predictions based on old data (by definition) and not the models, how do they just the predictions them to make trading decisions? * How can I invest in this hedge fund?
- ktamura 10y agoTheir dataset reeks of startup hustling. I just downloaded the training set [1] and plotted some of its descriptive statistics [2]. It looks that all features are uniform distributions and the response variable is Bernoulli coin-flipping. In layman's terms, you can't really come up with a good predictive model with this training set. I give them the benefit of the doubt that they wanted to have something in place to push the website live, but I cannot imagine any serious data scientist not noticing this. [1] http://datasets.numer.ai/57feb95/numerai_training_data.csv http://datasets.numer.ai/57feb95/numerai_training_data.csv [2] https://cl.ly/051G3Y2Z2O0W https://cl.ly/051G3Y2Z2O0W
- antognini 10y agoThe features are all encrypted using homomorphic encryption, which ends up mapping the overall distribuiton uniformly onto [0, 1]. If you play with the dataset, though, you'll find there are some significant correlations between the different features and you can make a model that does substantially better than chance.
- daveguy 10y agoI would like to see a comparison between random generated uniform features and these features. Can you fit the noise and come up with "statistically significant" predictions? If so, is anyone doing any better than can be done on random data? If no one is doing better than that, what is the likelihood that these aren't any better than monkey with a dartboard? I would love to see peer reviewed articles from numerai with some of their behind the scenes results.
- antognini 10y agoWhenever I make a model I always do cross-validation to make sure that I'm not overfitting. If we were just fitting random noise I would always see the performance of my test set being no better than chance (or worse).
- jamez1 10y agoThe key property that you should be testing for is dependence. It would be surprising if you could identify anything in the descriptive statistics, that would make this all child's play? I can't imagine any serious data scientist not knowing that.
- asdfologist 10y agoI don't see how Numerai can avoid the multiple comparisons problem [0]. If people submit thousands of random models, then some subset of them will do a fantastic job in predicting prices in historical simulations but do poorly under real market conditions. As long as the models are black boxes, there's likely no good way to distinguish them from noise. [0] https://en.wikipedia.org/wiki/Multiple_comparisons_problem https://en.wikipedia.org/wiki/Multiple_comparisons_problem
- dreamdu5t 10y agoThat's true only if stock market data is completely random and that there's no signal to predict on. That doesn't seem to be the case, considering hedge funds successfully use ML on historical data to capture alpha. You don't need a model to work forever to be successful.
- asdfologist 10y agoIt looks like you missed the point of my comment. I'm saying that Numerai won't be able to distinguish between zero alpha and positive alpha models, if all they're doing is running historical simulations on black boxes.
- blahblah3 10y agoThis can be mitigated by evaluating all the models on a hold-out test set (similar to what kaggle does and what was done in the netflix prize). The multiple comparisons problem is also mitigated by the fact that the models wont be completely random, there will likely be some positive correlation between them. edit: Also, by hoeffding's inequality the number of training examples needed for a given level of confidence is only logarithmic in the number of models (even assuming they are independent). See page 6 here: http://cs229.stanford.edu/notes/cs229-notes4.pdf http://cs229.stanford.edu/notes/cs229-notes4.pdf
- cjlars 10y agoThey would need a model built on top of the submitted models to weight or select which specific trading signals to act on. A plausible model might be walkforward testing on out of sample data, or as simple as trailing performance in live trading, or both. Still, with a large set of models, you won't be able to get around multiple comparisons problem entirely. That doesn't matter though because you only a need small amount of good signal to be profitable in the money management biz. The larger issue is whether the aggregate signal quality is high enough to be able to pay for development and trading costs.
- dschiptsov 10y agoSomeone got financing for a crowdsourcing of data mining of commercial data set. Very clever. It seems that only very few nerds are taking the theoretical impossibility of predicting the future seriously and we are missing the opportunity to get funding for some crappy models from the greater fools.
- Houshalter 10y agoI don't understand how this works. What exactly is being predicted? The data isn't a time series. The outputs are only binary. There are only 27 features. I don't understand how this represents market data at all. In fact they probably destroyed most of the information trying to convert the data to this format.
- markovbling 10y agoAwesome!
- dooglius 10y agoThis is fishy. The entire point of encrypted data is that one message cannot be distinguished from another without decryption (which would require the private key). In other words, the entire premise shouldn't work. This means that either bad encryption is being used (i.e. statistical information about the data is leaked) or the good results we see are just noise. Or, the whole thing is a scam to get funding: the best algorithms are planted, and the company just shuffles BTC between accounts it controls.
- jamez1 10y agoThe data is homomorphically encrypted, meaning you can do operations (such as add and subtract) on the ciphered message and it will also perform them on the underlying data.
- dooglius 10y agoYes, I realize that. The issue is that the result of any operation is also encrypted, which means that there should be no way to connect the target of the training data (encrypted or not) to the output of a function of encrypted data. Suppose the unencrypted data is (a,b,c) where a+b=c, and (x,y,z)=encrypt((a,b,c)). We have an addition function plus on encrypted data such that decrypt(plus(x,y))=a+b=c=decrypt(z), but it is not the case that plus(x,y)=z (at least, not if plus is computable in polynomial time, and assuming the encryption scheme is sound). If it were, we could statistically distinguish encrypt((a,b,c)) from encrypt((rand(),rand(),rand())) which would mean the encryption is not sound.
- evanpw 10y agoI read through the blog posts, and it seems like the encryption is order-preserving. It's designed to leak enough information to be useful in prediction, but not enough for users to trade on it independently.
- jamez1 10y agoThey could just be normalizing every data point to be between 0 and 1 by dividing by the range. That's a homomorphic encryption.. it passes your weird assumptions. I don't know why you're harping on about sound encryption, the point of this is to keep the statistical information intact in the cipher, without giving away the underlying market data.
- Dowwie 10y agoIs this an execution of "algorithm(model) as a service"?
- deleted 10y ago[deleted]
- deleted 10y ago[deleted]
- blazespin 10y agohttps://xkcd.com/1570/ https://xkcd.com/1570/
- martinko 10y agoHope they will have more luck than LTCM
- darkwebsss_111 10y agoIf you ever need the services of a professional for all your cyber/identity issues, then darkwebssolutions on gmail is the one you should consult. Or text +19193076946