3 ms·
questions running through my head: Is there some magic with the number 8? Why not 6? 11? Each of these 8 models were 7B models. What about using 80 x tinyllama
by SubiculumCode 3y ago
questions running through my head:
Is there some magic with the number 8? Why not 6? 11?
Each of these 8 models were 7B models. What about using 80 x tinyllama 1B models?
- lee101k 3y ago[dead]
- pushfoo 3y agoTL;DR: It looks like a convenient and explainable starting point for the proof-of-concept. Manually combining specialist variants is a known technique. This paper automates it with a router component which mixes 2 sub-models at any given time. Training 8 slight variants of a base seems safe and configurable compared to n > 16 specialists. The latter seems like the parts could interact unpredictably. Also, the memory usage seems predictable: it follows 2^m memory conventions by mixing 2 models at a time, so ~2x the memory is actively used at a time. I'm not up to date on the hardware implications, so it might not mean anything yet. It might one day if this approach works well enough to design around.