3 ms·
I think so too. But in general, it could also be due to other reasons: faster hardware, lower timeout for batched inference, optimizations like flash attention
by rasbt 3y ago
I think so too. But in general, it could also be due to other reasons: faster hardware, lower timeout for batched inference, optimizations like flash attention and flash attention 2, quantization, ...
I'd say that it's probably a mix of all of the above (incl some distillation).