7 ms·
we did a lot of our early experimentation with small networks. I don't think we went any smaller than 5 layers of 64 filters as we mentioned here: https://mediu
by vishvananda 6y ago
we did a lot of our early experimentation with small networks. I don't think we went any smaller than 5 layers of 64 filters as we mentioned here: https://medium.com/oracledevs/lessons-from-alpha-zero-part-5-performance-optimization-664b38dc509e https://medium.com/oracledevs/lessons-from-alpha-zero-part-5...
- jonath_laurent 6y agoAnd what were the results of these experiments? What error rate can you reach with the smallest network architecture you tried for example?
- vishvananda 6y agoUnfortunately I don't remember the exact numbers, but I think it was a couple percentage points worse than we were able to get with the large models.
- jonath_laurent 6y agoThis is interesting, thanks! Is there anything else you can tell me about the results of your experiments with small networks? I am really interested in this. For example: did you notice than increasing or decreasing network size required significant changes in other hyperparameters? Are small networks learning faster at the beginning of training before they start to plateau?