3 ms·
It seems like this can’t run all models, and needs custom ones trained from scratch: “ We introduce two new models: TurboSparse-Mistral-7B and TurboSparse-Mixtr
by russianGuy83829 2y ago
It seems like this can’t run all models, and needs custom ones trained from scratch: “ We introduce two new models: TurboSparse-Mistral-7B and TurboSparse-Mixtral-47B. These models are sparsified versions of Mistral and Mixtral […]. Notbly, our models are trained with just 150B tokens within just 0.1M dollars”.
It remains to be seen how good these custom models are.
- TOMDM 2y agoPaper for the sparcified mixtral models https://arxiv.org/abs/2406.05955 https://arxiv.org/abs/2406.05955
- deleted 2y ago[deleted]
- helloericsf 2y agoAgreed. custom models could be a hit or a miss.
- cpldcpu 2y agoIt's just continued pretraining to "heal" the damage caused by switching the activation functions and enforcing sparsity. Apparently they managed to recover original performance on standardized tests after continuing pretraining with the 150B tokens. There may be some more specialized knowledge lost that was not covered by their dataset.