4 ms·
Well, again, i'm not talking about a language model.... And, just because what I'm saying isn't especially likely to work, it's not obvious that it cannot. Ver
by danielmarkbruce 2mo ago
Well, again, i'm not talking about a language model....
And, just because what I'm saying isn't especially likely to work, it's not obvious that it cannot. Very large models are doing all manner of things that very smart people thought were not possible just 6 or 7 years ago.
- insanitybit 2mo agoIt's unclear what you are talking about then. Because the idea of training "ciphertext -> plaintext" for language models is absolutely bonkers, so what are you suggesting?
- danielmarkbruce 2mo agoI can't tell if you are serious at this point. I literally say, twice, that I'm not talking about a language model. And I also say the model would predict the key...not the plaintext.
- insanitybit 2mo agoAnd I'm asking you to describe the model.
- danielmarkbruce 2mo agoOuput: 128 logits. Input: maybe 10 samples of plaintext,ciphertext (using the same key), so maybe a 2560 length tensor. Loss function: binary cross entropy on the true key bits. Architecture: anyone's guess. If you were in a place to debate this, you would have known the above (or something similar) is what I was suggesting when i said train on plaintext, cipertext -> key, and you'd have some deep mathematical insight as to why no architecture known is likely to work. And you would also know I wouldn't be here talking to you about it if I really had a solid idea of an architecture that is likely to work.
- insanitybit 2mo agoI'm not debating you at all. I'm asking what the model looks like since you've stated (and I've agreed) that a language model wouldn't work. I think it would make sense to explain how a theoretical model could do better than SAT. Otherwise, is the idea here just "magic is possible"?
- danielmarkbruce 2mo agoYes, "magic is possible" if you defined "magic" as "very large models approximating functions in a way that people didn't think would work". Current SOTA language and vision models, or models used to predict protein shapes are magic by the standards of 2016. As for why could it be better than a SAT? Why couldn't it be? Models are better than deterministic, logically written software for lots of situations. You can create infinite training data for this problem. The number of humans that work on encryption is tiny. The idea that because humans haven't figured out how to break some encryption schemes it can't be done is kind of absurd.
- insanitybit 2mo agoSure, that seems reasonable enough. I'm pretty skeptical that it will happen, but it's not like it's impossible.
- SideQuark 2mo agoHave you ever built a NN model? Have you ever broken a crypto system, even a small one?
- danielmarkbruce 2mo agoyes and yes. You can build and train a model in about 15 lines of pytorch. And you can build and break your own 8 bit xor cipher in about 10 lines of python. Hacker news is full of software engineers. You are unlikely to find one that hasn't built a model using pytorch these days, and an xor cipher is a common university lab exercise.
- SideQuark 2mo agoWhen you conflate an 8-bit xor cipher with AES and a 10 line python NN with the complexity of trying to break AES with a NN by having "the model ... predict the key", I can see how you think such a thing is possible. It's nonsense. For any NN to learn, the function it is approximating must structured enough to admit small set of parameters (i.e., not exponential), otherwise it will take exponentially many nodes to do anything. AES is not differentiable, as are all secure hash and encryption functions. This is a basic test that is used to attack everything. So you'd get a NN that must be big enough to simply memorize all plaintext, key, output triplets, which is simply a lookup table. With around 2^768 nodes. Good luck.
- danielmarkbruce 2mo agoLol, have you heard of "moving the goalposts"? Could you possibly be more insincere?