3 ms·
So there's no performance gain for quantization enabled by the transformer architecture? It seems very strange that quantization works so well since in most of
by grungegun 3y ago
So there's no performance gain for quantization enabled by the transformer architecture? It seems very strange that quantization works so well since in most of my experiments, the internal model weights of mlps look random.