3 ms·
RLHF is a popular candidate, but the focus is more on "helpfulness" and "safety" -- I don't think it necessarily improves LLMs on reasoning benchmarks
by rasbt 3y ago
RLHF is a popular candidate, but the focus is more on "helpfulness" and "safety" -- I don't think it necessarily improves LLMs on reasoning benchmarks
- behnamoh 3y agoif anything, RLHF makes the model dumber, not smarter.
- rasbt 3y agoI think it could potentially make the model smarter, but it's up to how you collect the data to train the reward models. Currently, companies & papers that use RLHF focus on "safety" rankings, for example. But you could potentially collect labels "smartness" or "correctness" instead and train the the reward model one these. (And then use that reward model to finetune the LLM you want to improve.)