4 ms·
For short context tasks looks like it's slightly stronger than Llama 7B and slightly weaker than Mistral 7B. Really impressive showing for a completely new arch
by kcorbitt 3y ago
For short context tasks looks like it's slightly stronger than Llama 7B and slightly weaker than Mistral 7B. Really impressive showing for a completely new architecture. I've also heard that it was trained on far fewer tokens than Mistral, so likely still room to grow.
Overall incredibly impressive work from the team at Together!
- tempusalaria 3y agoDid they disclose the training compute/token count?
- stellaathena 3y agoNope :(