3 ms·
It's a disadvantage of current SOTA models: they are easy to train, but they must be large wasting lots of weights in order to generalize well. Maybe another ar
by novaRom 3y ago
It's a disadvantage of current SOTA models: they are easy to train, but they must be large wasting lots of weights in order to generalize well. Maybe another architecture, transformer's successor will be more economical - having less weights with more skills and knowledge.
- two_in_one 3y agoI think I've seen somewhere years ago an article claiming size is needed for training. Which means probably that after training model can be optimized to minimize the size. Purging? However, smaller model cannot be trained that well. From my experience with image processing the bigger the better, and, they all have their limits. Nothing new here.