3 ms·
On the limitation side: Do you think this would scale to larger transformer models with more parameters per layer? How would this work with MOE models or spar
by gkapur 5mo ago
On the limitation side:
Do you think this would scale to larger transformer models with more parameters per layer?
How would this work with MOE models or sparse models?