4 ms·
What kind of accuracy did you get with the transfer learning attempts?
by bigfish24 9y ago
What kind of accuracy did you get with the transfer learning attempts?
- timanglade 9y agoWell for a while I was lulled into complacency because the retrained networks would indicate 98%+ accuracy, but really that was just an artifact of my 49:1 nothotdog:hotdog image imbalance. When I started weighing proportionately, a lot of networks were measurably lower, although it’s obviously possible to get Inception of Vgg back to a “true” 98% accuracy given enough training time. That would have beat what I ended up shipping, but the problem of course was the size of those networks. So really, if we’re comparing apples to apples, I’ll say none of the “small”, mobile-friendly neural nets (e.g. SqueezeNet, MobileNet) I tried to retrain did anywhere near as well as my DeepDog network trained from scratch. The training runs were really erratic and never really reached any sort of upper bound asymptotically as they should. I think this has to do with the fact that these very small networks contain data about a lot of ImageNet classes, and it’s very hard to tune what they should retain vs. what they should forget, so picking your learning rate (and possibly adjusting it on the fly) ends up being very critical. It’s like doing neurosurgery on a mouse vs. a human I guess — the brain is much smaller, but the blade says the same size :-/
- bigfish24 9y agoVery interesting! If you were to make a v2, would you adjust the 49:1 imbalance and add more hot dog images?
- timanglade 9y agoI’m not sure, I think I would maybe break classes into multiple labels, but that becomes even more finicky to train. At the end of the day, there are many more things that are not hotdogs, than things that are hotdogs, so you do have to provide more examples of the not hotdogs to train something from scratch properly — I don’t see a way around it. Honestly I think the biggest gains would be to go back to a beefier, pre-trained architecture like Inception, and see if I can quantize it to a size that’s manageable, especially if paired with CoreML on device. You’d get the accuracy that comes from big models, but in a package that runs well on mobile.