3 ms·
ChatGPT was also trained with reinforcement learning to produce outputs human raters preferred
by fiso64 4y ago
ChatGPT was also trained with reinforcement learning to produce outputs human raters preferred
- m00x 4y agoIt's definitely getting closer in this case, but still chatGPT still look for any truth in its response.