2 ms·
" a judge agent then attempts each task to verify that it is actually solvable " I understand you need to verify the goal is achievable. But if the judge agent
by smurf9852 2mo ago
" a judge agent then attempts each task to verify that it is actually solvable "
I understand you need to verify the goal is achievable. But if the judge agent has the same goal as the training agent (solve), and both are of the same model, then aren't the judge and the training agent doing the exact same thing? What is the point then? Can someone explain this to me.
- deleted 2mo ago[deleted]
- alightsoul 2mo agoIt works because models are not deterministic so you can prompt two instances of the same model to be adversaries of each other