3 ms·
Hi, one of the lead authors for this work. We recommend using Bfloat16 (not fp16), quantization for small models can really hurt performance!
by mluo 2y ago
Hi, one of the lead authors for this work.
We recommend using Bfloat16 (not fp16), quantization for small models can really hurt performance!
- CamperBob2 2y agoHave you compared it to the 1.58 bit dynamic quant model based on the original R1 (i.e., not a distillation)? Whatever unsloth did, it doesn't seem to be giving up much reasoning performance over the full Q8 version.
- mluo 2y agoIt's simply bc the model is small (1.5B), making it sensitive to weight perturbations
- simonw 2y agoIs there a GGUF version of your model anywhere that you recommend? I'm on a Mac.
- mluo 2y agoThink there are some people who made GGUFs as branches of our model, try it out! https://huggingface.co/models?other=base_model:quantized:agentica-org/DeepScaleR-1.5B-Preview https://huggingface.co/models?other=base_model:quantized:age...
- newman314 2y agoIs there a MLX version that can be added to the fullmoon iOS app?