4 ms·
For anything past 8k context size We are talking about over 10x reduction in GPU time for inferencing tokens and for training too Aka it’s cheaper and faster
by pico_creator 3y ago
For anything past 8k context size
We are talking about over 10x reduction in GPU time for inferencing tokens and for training too
Aka it’s cheaper and faster
Alignment is frankly IMO purely a dataset design and training issue. And has nothing to do with the model