2 ms·
Plausible, although I don't see anything about reference solutions in the ExploitGym paper or github. Doesn't mean they don't exist, but it's not obvious to me
by Imnimo 2mo ago
Plausible, although I don't see anything about reference solutions in the ExploitGym paper or github. Doesn't mean they don't exist, but it's not obvious to me that we should expect to find these on HuggingFace.
- ollin 2mo agoThe ExploitGym paper evaluated several frontier models on the bench and reported that "Different models find different exploits" [1], so it seems most plausible that the "test solutions directly from Hugging Face’s production database" [2] which GPT-internal found were authored by Mythos (or some other LLM with complementary strengths), and placed in some internal HF repository when creating the ExploitGym paper/leaderboard. [1] https://www.cybergym.io/exploitgym/#:~:text=Different%20models%20find%20different%20exploits https://www.cybergym.io/exploitgym/#:~:text=Different%20mode... [2] https://openai.com/index/hugging-face-model-evaluation-security-incident/#:~:text=test%20solutions%20directly%20from%20Hugging%20Face%E2%80%99s%20production%20database https://openai.com/index/hugging-face-model-evaluation-secur...