3 ms·
TL;DR: It looks like a convenient and explainable starting point for the proof-of-concept. Manually combining specialist variants is a known technique. This pa
by pushfoo 3y ago
TL;DR: It looks like a convenient and explainable starting point for the proof-of-concept.
Manually combining specialist variants is a known technique. This paper automates it with a router component which mixes 2 sub-models at any given time. Training 8 slight variants of a base seems safe and configurable compared to n > 16 specialists. The latter seems like the parts could interact unpredictably.
Also, the memory usage seems predictable: it follows 2^m memory conventions by mixing 2 models at a time, so ~2x the memory is actively used at a time. I'm not up to date on the hardware implications, so it might not mean anything yet. It might one day if this approach works well enough to design around.