Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mluo
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
mluo
2y ago
Check out one of my prior work: https://stylus-diffusion.github.io/ This work scales up selection/routing over many models/LoRAs
2.
▲
by
mluo
2y ago
For quantization, very big impact for small models, can drop at much as 10% on AIME. Our model does best on bfloat16 ;) Come checkout our repo at: https://github.com/agentica-project/deepscaler
3.
▲
by
mluo
2y ago
It's simply bc the model is small (1.5B), making it sensitive to weight perturbations
4.
▲
by
mluo
2y ago
Think there are some people who made GGUFs as branches of our model, try it out! https://huggingface.co/models?other=base_model:quantized:age...
5.
▲
by
mluo
2y ago
Nice, very glad to see it works! Small models are very sensitive to the dtype :(
6.
▲
by
mluo
2y ago
Try bfloat16! We have a bug where the model was saved as fp32.
7.
▲
by
mluo
2y ago
We beat O1-preview and even many other 7B models over many math benchmarks, which was TEST set (not in training set at all). If you want to make the model fully generalist, feel free to train it over coding datasets (such as RL with passing
8.
▲
by
mluo
2y ago
One of the authors here.... This is not a Chinese model, btw I'm American
9.
▲
by
mluo
2y ago
Hi, one of the lead authors for this work. We recommend using Bfloat16 (not fp16), quantization for small models can really hurt performance!
10.
▲
DeepScaleR: Surpassing O1-Preview with a 1.5B Model by Scaling RL
(pretty-radio-b75.notion.site)
19 points
by
mluo
2y ago
|
0 comments
11.
▲
by
mluo
3y ago
Alpaca Llama Vicuna -> Gorilla Chad move
12.
▲
by
mluo
4y ago
With inflation in mind, wouldn't there a larger gap between Ray's sort and the previous WR holder?