3 ms·
Why would you do this?
by Demiurge 2mo ago
Why would you do this?
- teravor 2mo agoa specialist model sufficiently post-trained can outperform a frontier model while being dirt cheap. what we did was distill GLM 5.2 into a 27B model on SQL and then post-train it with RL afterward. the result outperformed even Fable on that one task. the distillation step is just good sense in this workflow, to bootstrap a smaller model to the utmost you can before actually doing RL.