3 ms·
I just tried Mistral MoE 8x7B model and it works a bit faster than llama-2-70B but it looks it has almost the same skills. In fact, all latest models of 13B-70B
by novaRom 3y ago
I just tried Mistral MoE 8x7B model and it works a bit faster than llama-2-70B but it looks it has almost the same skills. In fact, all latest models of 13B-70B size are quite similar. Could it be large part of their training data is the same?
- cyanydeez 3y agothere might be a laws of average thing going on? Central Limit Theorem? there's a finite number of relevant language tokens. the trick of current LLM is finding what's basically the center of a vast series of probability.