4 ms·
Multi-Query Attention, used here, should make 40B inference viable on systems where even 33B LLaMA with Multi-Head is basically unusable, so sometimes improveme
by airgapstopgap 3y ago
Multi-Query Attention, used here, should make 40B inference viable on systems where even 33B LLaMA with Multi-Head is basically unusable, so sometimes improvements still come from software optimization (it's no free lunch though).