4 ms·
I didn't quite get why RL was used instead of just collecting same data from labelers and either prompt-tuning, p-tuning or fine-tuning on that.
by rllearneratwork 5y ago
I didn't quite get why RL was used instead of just collecting same data from labelers and either prompt-tuning, p-tuning or fine-tuning on that.
- sanxiyn 5y agoRead the paper. They did finetune from labeled data and set it as baseline. RL outperformed baseline.