3 ms·
“On an ordinary coding prompt, the J-space of a model trained to sabotage code contains “fake,” “fraud,” “secretly,” and “deliberately” at the start of its resp
by SequoiaHope 3mo ago
“On an ordinary coding prompt, the J-space of a model trained to sabotage code contains “fake,” “fraud,” “secretly,” and “deliberately” at the start of its response.”
I would like to know more about their model trained to sabotage code…
- tough 3mo agohttps://arxiv.org/pdf/2511.18397 https://arxiv.org/pdf/2511.18397
- SequoiaHope 3mo agothank you!