3 ms·
It still does a much better job at translation than llama 2 70b even, at 6.7b params
by ronyfadel 3y ago
It still does a much better job at translation than llama 2 70b even, at 6.7b params
- two_in_one 3y agoIf it's MOE that may explain why it's faster and better...
- yumraj 3y agoMOE?
- sarthaksrinivas 3y agoMixture of Experts Model - https://en.wikipedia.org/wiki/Mixture_of_experts https://en.wikipedia.org/wiki/Mixture_of_experts