7 ms·
Sorry if this is a dumb question. Can someone explain why it’s called 8x7B(56B) but it has only 46.7B params? and it uses 12.9B params per token generation but
by chandureddyvari 3y ago
Sorry if this is a dumb question. Can someone explain why it’s called 8x7B(56B) but it has only 46.7B params? and it uses 12.9B params per token generation but there are 2 experts(2x7B) chosen by a 2B model? I’m finding it difficult to wrap my head around this.
- pilotneko 3y agoI haven’t looked at the structure carefully, but It’s hard to guess there are shared layers between models. Likely the input layers for sure, since there is no need to tokenize separately for each model (unless different models have specialized vocabulary).
- lordswork 3y agoThis is my understanding as well. Also includes the parameters of the expert-routing gating network.
- brrrrrm 3y agomixture of experts gates on the feed forward network only. the shared weights are the KQV projections for the attention mechanism of each layer.
- shekhar101 3y agoExplanation from Andrej karpathy makes sense on why: ''' "8x7B" name is a bit misleading because it is not all 7B params that are being 8x'd, only the FeedForward blocks in the Transformer are 8x'd, everything else stays the same. Hence also why total number of params is not 56B but only 46.7B. '''