3 ms·
Maybe an adversarial approach was used in training these models in the first place?
by arrow7000 4y ago
Maybe an adversarial approach was used in training these models in the first place?
- sharemywin 4y agoIt was they were' trained using reinforcement learning with human feedback to create the critic.
- jedberg 4y agoI hadn't thought about human feedback being an adversarial system, but I guess that makes sense, since it's basically a classifier saying "you got this wrong".