3 ms·
I don't think overfitting will be an issue (at all!) on a 65TB dataset. A bigger CNN model should be more effective at this task than a DBN, (almost) regardless
by kastnerkyle 11y ago
I don't think overfitting will be an issue (at all!) on a 65TB dataset. A bigger CNN model should be more effective at this task than a DBN, (almost) regardless of the features added. If we can use convolutional generative models (from raw noise -> images) to make this kind of stuff [1][2] I see no reason why it shouldn't be super effective to classify with a CNN.
The equivalent DBN-type generative model for CIFAR10 is pretty far behind [3], but not terrible. It is believable that a DBN could do well on this task, but it would be very, very surprising if a DBN beats a well trained CNN of the appropriate capacity in anything related to classification.
All that said, some of the other comments have highlighted larger potential issues than model choice here - achieving up to 98% accuracy using single pixels, random forest, etc seems to point to potential issues in the dataset that will block any kind of model evaluation or further research.
I would look for data leakage, and reconsider CNNs in your future work - especially something using larger patches and VGG style features. With respect to your comment on scalability - if the model converges and can't really learn more (small/limited capacity network) processing more training data is just a waste of time. Convergence is not really epoch based, but update based, especially in big datasets like this.
A small network will be very fast to apply but if that is the goal there are lots of papers on model approximation - a big ensemble of networks that are engineered and model approximated to fit your compute budget might work better than limiting the network capacity initially.
All said - does it really matter much if you get 95% vs. 97% vs. 98% vs. 99.99%? Aren't there meta techniques like CRF to resolve occasional blips in model prediction for neighboring patches anyways? Maybe a linear model with decent features or random forest + follow-on cleanup will work better with the constraints necessary.
I am all about neural networks for most things, but if you have a serious computational constraint linear models (or random forests) are stupid fast [4], and pretty good on many tasks. Adding on a "meta model" to resolve anomalous errors with this could be good enough for your task. Just something to consider.
[1] https://twitter.com/AlecRad/status/645830923150299136 https://twitter.com/AlecRad/status/645830923150299136
[2] http://soumith.ch/eyescream/ http://soumith.ch/eyescream/
[3] http://www.icml-2011.org/papers/591_icmlpaper.pdf http://www.icml-2011.org/papers/591_icmlpaper.pdf
[4] https://github.com/ajtulloch/sklearn-compiledtrees/ https://github.com/ajtulloch/sklearn-compiledtrees/