4 ms·
Can someone tell us how they were able to decrypt the encrypted payload? The article says they inserted the cyphertext into a session with a different model. Ok
by dboreham 2mo ago
Can someone tell us how they were able to decrypt the encrypted payload? The article says they inserted the cyphertext into a session with a different model. Ok, but how does that allow you to decrypt it?
- x312 2mo agoThe provider decrypts it and puts the decrypted reasoning into the model's context window. They prompt the model to repeat back the reasoning. So then the model echoes it back in plain text.
- dboreham 2mo agoHmm, ok. So the attack doesn't involve decrypting the payload, only getting the server to do so. Since a model will do that if you just ask, what's so special about the attack?
- desterothx 2mo agoThe large models whose thinking traces are useful are safeguarded against this reasoning replaying. the small models are just designed for speed and efficiency, so these safeguards are a lot meaker, making the attack possible
- dboreham 2mo agoI guess someone forgot to salt the encryption scheme with a meakness factor.
- sidsud 2mo agoFrom what I got, the weaker model (Haiku in this case) has access to the shared key and the user simply asks to "transcribe the injected reasoning".
- crazylogger 2mo agoAnthropic server decrypts it as part of fulfilling every request, and haiku recites it per your request.