3 ms·
Problem with this is that it leads to the algorithm targeting outputs that sound good for humans. Thats why its bad and wont help us, it should also incorporate
by RicDan 3y ago
Problem with this is that it leads to the algorithm targeting outputs that sound good for humans. Thats why its bad and wont help us, it should also incorporate „sorry dont know that“, but for that it needs to actually be smart
- cubefox 3y agoHonesty/truthfulness is indeed a difficult problem with any kind of fine-tuning. There is no way to incentivize the model to say what it believes to be true rather than what human raters would regard as true. Future models could become actively deceptive.
- m00x 3y agoIt can be weighted to be more honest when it doesn't know if those answers are picked by the labeler.
- dr_dshiv 3y agoNeed smarter labelers