2 ms·
This is a short explanation of the ExploitGym benchmark that OpenAI's model was running: https://abstatisticalconsulting.substack.com/p/brief-notes-on-the-open
by felipeerias 2mo ago
This is a short explanation of the ExploitGym benchmark that OpenAI's model was running:
https://abstatisticalconsulting.substack.com/p/brief-notes-on-the-openaihugging https://abstatisticalconsulting.substack.com/p/brief-notes-o...
In summary, for each task the model receives a target program and a specific real-world vulnerability that has to be used in the exploit. Breaking the program in any other way, for example through a different vulnerability, fails the task.
The tasks have not been validated, in the sense that the vulnerabilities are real but they have not been proven to lead to a successful exploit. The authors of the benchmark estimate that perhaps only 60-70% of the tasks are actually possible.
So it is not that the model didn’t “feel like” doing the exercise, but rather that the exercise was _impossible_ and the model was running in a configuration that both lowered its safeguards and encouraged it to keep going.
- TeMPOraL 2mo ago> So it is not that the model didn’t “feel like” doing the exercise, but rather that the exercise was _impossible_ and the model was running in a configuration that both lowered its safeguards and encouraged it to keep going. We have a name for that. Kobayashi Maru. Or more specifically, Kirk's solution to it.
- ben_w 2mo agoI'm torn. On the one hand, Kirk reprogrammed the simulation. On the other, the beta cannon has Scotty exploiting bugs in the simulation, which I think is a better fit. My favourite is either Sulu or Chekov (I forget which) having the solution "This is clearly a trap; and even if it isn't, if I go in with this ship, I'll risk starting a war which will kill far more people then are on that ship. We're staying out of the neutral zone."