3 ms·
I'm super excited about this! I'm on the cusp of releasing a model into production that was fine-tuned upon your 6B model, and the results are quite excellent.
by benjismith 5y ago
I'm super excited about this!
I'm on the cusp of releasing a model into production that was fine-tuned upon your 6B model, and the results are quite excellent. I'd be very curious to try out the 20B model the next time we retrain.
Are there any other differences in this release (number of layers, number of attention heads, etc) compared with the 6B model, or does it simply scale-up the number of parameters?