3 ms·
Hah, I wrote this article ages ago when we were experimenting with a data science blog. Surprised to see it pop up here. I think most of it still holds true bu
by micro_cam 11y ago
Hah, I wrote this article ages ago when we were experimenting with a data science blog. Surprised to see it pop up here.
I think most of it still holds true but we did do some benchmarks on public data and the implementation is competitive with or faster then other implementations including scikit-learn's cython implementation which wins a lot of benchmarks.
Code and some benchmark results in the README are here:
https://github.com/ryanbressler/CloudForest https://github.com/ryanbressler/CloudForest
- rgbrgb 11y agoI submitted it because I'm playing with CloudForest right now. I really appreciate your work, it's fun to use. I'm co-founder of a digital real estate brokerage[1] and as a weekend hack, I'm using CloudForest to predict sale prices for residential real estate using data from the MLS. A lot of the predictions are really good with a few that are terrible (like double the list price, which is a feature). Any way to get a certainty value for each prediction? Perhaps I could look at the variance of values from each tree? I'm kind of a noob with decision trees so I'm probably totally off base but maybe that question makes sense. :) [1]: https://www.openlistings.co/ https://www.openlistings.co/ (YC W15)
- pigscantfly 11y agoYes, you should be able to estimate a confidence interval by looking at the distribution of votes from each tree. [1] http://arxiv.org/pdf/1311.4555.pdf http://arxiv.org/pdf/1311.4555.pdf