4 ms·
It seems like OpenAI's Human Feedback research (in collaboration with DeepMind) is targeted at this sort of thing. They try to use human feedback to create more
by colah3 9y ago
It seems like OpenAI's Human Feedback research (in collaboration with DeepMind) is targeted at this sort of thing. They try to use human feedback to create more nuanced and aligned objectives.
https://blog.openai.com/deep-reinforcement-learning-from-human-preferences/ https://blog.openai.com/deep-reinforcement-learning-from-hum...
- deleted 9y ago[deleted]