3 ms·
The main thing that led me to believe it's not trained directly on benchmarks is the fact that this model - a direct finetune of Mistral, improved on them. Con
by The_Contrarian 3y ago
The main thing that led me to believe it's not trained directly on benchmarks is the fact that this model - a direct finetune of Mistral, improved on them.
Consider the following thought experiment. Let's say we train a glorified benchmark table. We train it extensively on benchmarks from a number of different providers. MMLU, Hellaswag, Winogrande, etc.
Such a model will do well on benchmarks, but its reasoning ability, obviously, will be severely below what we'd expect from those benchmarks.
What happens when we tune it on a dataset like Orca? Will we observe improvements on the benchmarks? Well no - if anything, we might even expect some regression as the model doesn't have sufficient reasoning ability to handle these directly, just the answers it was fed during pretraining.
The fact that we don't observe this with Mistral, and that finetunes that are designed to improve reasoning ability lead to an increase in the corresponding benchmarks as we'd expect, leads me to believe that Mistral's performance is legitimate.
Even apart from that - it's worth checking out StableLM's recent 3B release. While it didn't get as much attention due to literally releasing the day after Mistral and they didn't promote it until afterward, StableLM 3B boasts some significant gains in comparison to its predecessor as well due to training for significantly longer: https://stability.wandb.io/stability-llm/stable-lm/reports/StableLM-3B-4E1T--VmlldzoyMjU4?accessToken=u3zujipenkx5g7rtcj9qojjgxpconyjktjkli2po09nffrffdhhchq045vp0wyfo https://stability.wandb.io/stability-llm/stable-lm/reports/S... and these gains are likely to be more substantial at larger parameter sizes.
There's still a lot to be gained from training models for longer, and I can fully believe that Mistral's level of performance is possible in a 7B model.