3 ms·
First, this is not an open source / weight release. Second, it has the problem of non-stoping response.
by RandyOrion 2y ago
First, this is not an open source / weight release.
Second, it has the problem of non-stoping response.
- inciampati 2y agoWhat's the best technique to train the model to stop responding? A bit of fine tuning on texts with EOS markers?
- RandyOrion 2y agoI didn't see many papers on solving this problem. I see non-stop response as a generalization problem because normally every training sample is not of infinite length. Targeted supervised fine-tuning should work, as long as you have enough samples. However, supervised fine-tuning is not good for generalization.