3 ms·
The interesting part is that distillations based on reinforcement learning based models are performing so well. That brings the cost down dramatically to do cer
by physicsguy 2y ago
The interesting part is that distillations based on reinforcement learning based models are performing so well. That brings the cost down dramatically to do certain tasks.
- azinman2 2y agoI thought the distillations were SFT only?
- physicsguy 2y agoThey're SFT on the chain of thought output of R1