6 ms·
LLMs won't break symmetric crypto
- amingilani 2mo ago> They’re time- and battle-tested All conjectures are until someone with the time and energy proves or disproves them.
- sghiassy 2mo agoA next-word-in-the-sentence prediction engine can’t predict the factor of two insanely large prime numbers… tell me more
- dadrian 2mo agoRSA is asymmetric crypto. This article is about symmetric cryptography. I expect LLMs will advance state of the art in factoring algorithms, considerably.
- what 2mo agoWhy?
- sghiassy 2mo agoThank you I guess I only know asymmetric cryptography. I should learn more about symmetric… Anyone care to boil it down for me :) Edit: Isn’t this just advanced static analysis of any code base?
- fluoridation 2mo agoVery, very briefly, most symmetric algorithms are block ciphers, meaning that their input are blocks of a fixed length in bits (plus a key), and their output is another block of the same length. Ideally, a block cipher with its key produces a random permutation of the input space into the output space, thus diluting the information and dramatically increasing (ideally maximizing) the entropy; what that means is that whether the input is just zeroes and ones in ASCII or fully random, after encryption it should be indistinguishable.
- sghiassy 2mo agoThank you I wish I knew more in this domain. It almost sounds like hashing with a salt
- fluoridation 2mo agoIt pretty much is, except it's reversible. At the block level it meets the cascading requirement, and you can set it up to expand the output arbitrarily by padding the input with zeroes (thus also turning it into a PRNG).
- retrac 2mo agoA symmetric cipher is: ciphertext = data XOR key. XOR is reversible: plaintext = ciphertext XOR key. If the key is a set of truly random numbers the same size as the ciphertext, then this is a one-time pad, and it is truly secure in the information theory sense. Nothing other than knowing the original randomly selected key values can decode the ciphertext. But of course, it's hard to come up with terabytes of random numbers at the drop of a hat, and to share them securely with the other party. So symmetric ciphers use pseudo-random generation techniques, to iterate through many pseudo-random keys based on one original key. With PRNGs the "randomness" may have patterns and that is the opening for a break in the crypto.
- volkercraig 2mo agoThere is already a mathematically secure algorithm for securing a message: One Time Pad. The problem is that OTP requires that the length of the key and the length of message must be the same, which is inconvenient for large amounts of data. So the solution is to find algos that let you use a smaller key, but the side effect is that by pigeonhole principle, your keyspace is smaller than the message space, so it MUST be insecure. The trick is to make it so that it's only insecure enough that it's infeasible to break.
- tptacek 2mo agoIt's inconvenient for any amount of data, because it essentially begs the question; if you can securely transmit N bytes of key pad to a counterparty, just use that mechanism to transmit N bytes of plaintext instead.
- inigyou 2mo agoIt has the advantage that the key can be sent before the message is known. Think military battlefield. Your commander goes out to war with a CD, and then he can transmit messages like "we encountered the enemy". It would do no good to transmit "we encountered the enemy" before the war started.
- catlifeonmars 2mo agoPerhaps, but it’s still trivially easy to increase the difficulty of factorization problems on classical computers, We need a machine that can run Shor’s algorithm before integer factorization is practical and we’re still a long way out f M that.
- OJFord 2mo agoI think the thing most of us missed in dismissing GPT 2-3 as 'next word in sentence predictors' was that recursively this allows something resembling thinking, 'reasoning'. LLMs are capable not just of calculating the most likely next word from a prompt according to a corpus of training text, but of doing so & feeding back into themselves, the most likely word now based not only on the corpus but on the basic prediction, a second (nth) stage of thought. Yes it's all still token prediction, but it's predicting conversation between let's say not experts but capable speakers with all the information at hand. Undergraduates if you like. And such conversation can yield real results.
- sghiassy 2mo agoI’m with ya I’ve even heard arguments that prediction is consciousness. But using a Language-Model to break cryptography is still a stretch for me. From the little I know, cryptography uses information theory to make sure that reversing the equation (aka finding the passowrd) is predictably impossible, given current compute standards for the foreseeable future (disregard quantum computer here though :) they’re not LLMs)
- PlasmaPower 2mo agoThe oversight in your thinking is that we have no proofs about how much computation is needed to break cryptography. For all we know, it could be possible to break all modern cryptosystems in under a second on a computer from a decade ago with the right algorithms. This is how cryptography has been broken in the past: not just advances in the amount of compute we can do, but exponential speedups in the algorithms to break them. While I agree with the author of this post that modern cryptosystems are very secure and LLMs are not currently near breaking them, I don't think it's unreasonable to consider that if LLMs continue to get exponentially smarter they may make strides in cryptanalysis that we had never considered and break cryptography in unexpected ways. After all, many past cryptography breaks have come from previously unknown methods of cryptanalysis.
- sghiassy 2mo ago
- whateveracct 2mo agookay so silicon valley won't happen all the way
- modeless 2mo agoI don't really find the "because it's difficult" arguments convincing at all. Especially the one claiming it's hard because it requires designing and running a large number of tests and reasoning about the results of each one. That kind of tedious grinding is exactly where LLMs should shine vs humans! The only convincing argument here is that these things are battle tested (literally in most cases I would guess), with tons of research that never gets published because it's unsuccessful. A whole lot of human effort has gone into trying to break these things. A lot more than went into any of the math problems AI has solved so far. It's going to take a while before LLMs can equal and surpass that amount of human effort. And they might have to surpass it by many, many times to actually break these, if it is even possible, which is not certain.
- dboreham 2mo agoI read it as "because there are no viable attacks", which is...fightin' talk I suppose. What I have seen LLMs do recently is find what turned out to be very basic bugs in encryption and ZK libraries that for some reason humans never saw. In those cases it wasn't that the encryption algorithms were broken per se, but the the implementation was. This alone seems very worthwhile.
- modeless 2mo agoAgreed, we have probably seen only the tip of the iceberg on that. I wouldn't want to be holding niche crypto coins right now.
- cyberax 2mo agoThere are very few computer-era symmetric ciphers that were truly broken. RC4 is probably the worst example. There are no reasonable attacks even on the good old DES. And by "reasonable" I mean attacks that would bring down the complexity to a practical level if the DES key size were to be extended to something like 128 bits. We can brute-force DES keys trivially, but that's not a fault of the cipher per se.
- tptacek 2mo ago
- zkmon 2mo agoCryptographic systems are based on 1) mathematical impossibility of reversing some integer/mod calculation, 2) time required for a brute force attack, 3) correctness of algorithms and code used in implementations. The last part (algorithms and code) is where LLMs have a chance. The first one is not similar to the mathematical breakthroughs LLMs are making recently. There is a loss of information in mods and integer computations making them one-way. The second one requires simply increasing bit-length to match the increased computer power.
- stingraycharles 2mo agoYeah, I wouldn’t say with certainty that LLMs will never break any symmetrical crypto algorithm. It will certainly require a lot of effort, but so does solving some hard math challenges and it has been proven successful in that in the past. Most likely outcome will be that a security researcher is able to break one with assistance of / in collaboration with an LLM.
- tptacek 2mo agoSymmetric cryptography isn't based on complicated math the way asymmetric cryptography it is. The right way to think about symmetric cryptography is that the core hard problem is simply making PLAINTEXT XOR KEY work, efficiently, with a key that repeats.
- stingraycharles 2mo agoIsn’t another aspect of it that it’s sufficiently random / unrecognizable, for example? I’m very much aware of the differences between symmetric and asymmetric encryption, and realize that symmetric encryption is much simpler, but I figure that if there are weaknesses to be found in algorithms such as md5, then surely there are also potential weaknesses in symmetric encryption algorithms? Now I’m not saying that this would be the case for battle tested algorithms like AES. But is there any particular reason why this whole category could not possibly have weaknesses?
- 2mo ago
- arberx 2mo agoLLMs will accelerate math research, increasing understanding in areas like quantum which will eventually lead to breakthroughs that will break most standard asymmetric encryption algorithms with the side effect of breaking crypto
- deleted 2mo ago[deleted]
- tiahura 2mo agoArvin Krishna says 4 years
- tptacek 2mo agoRight, maybe, but Aumasson's whole point is that this prediction doesn't apply to AES, SHA2, BLAKE2, &c.
- TheDong 2mo agoWhat makes your prediction more likely than: "LLMs will accelerate math research, allowing us to prove that meaningfully sized quantum computers are impossible and crypto is secure. Modern cryptographic algorithms remains unbroken until the last human is turned into a paperclip in the year 2430"
- deleted 2mo ago[deleted]
- SideQuark 2mo agoThere are plenty of algorithms that quantum algorithms give no benefit to. Quantum is NOT a free for all making everything faster. It’s only faster at two very,very restricted things: hidden subgroup problem (which is used for RSA, but most problems do not rely on HSP), and Grover search, which reduces a specific search problem from O(N) to O(sqrtN).
- dsp_person 2mo agoWhat about checking crypto libraries for gaps like the coldcard situation of RNG code is correct but not in the release build somehow?
- krupan 2mo agoThat was such a stupid coding/code review/testing mistake. Finding it is not that impressive at all. It's nothing like finding a flaw in AES
- danielmarkbruce 2mo agoThis is kind of a stupid argument. How about make a slightly stronger claim like "models won't break symmetric crypto" ? I mean, language models aren't even trained to break symmetric crypto. There is not good reason to think they will. It seems possible to train a large model to do it though.
- catlifeonmars 2mo agoAgreed that many of the articles claims are a bit weak. One point is reasonably strong though: symmetric crypto may not be breakable (battle tested).
- danielmarkbruce 2mo agoIt probably isn't. But if you laid 20-1 I'd bet a large model will break an industry used standard within 10 years. That's a loose framing of a bet, but I think you get my point, even if you think my numbers suggest too much optimism.
- Monarch909 2mo ago[dead]
- tptacek 2mo agoTo train a large model to do what? Break AES? How would that work?
- danielmarkbruce 2mo agoTrain on plaintext, ciphertext -> key.
- tptacek 2mo agoLLMs aren't literally science fiction.
- 2mo ago
- random_mutex 2mo agoLLMs by themselves no, people with LLMS yes
- coderatlarge 2mo agobreaking some of these systems that humanity has been banging on for decades would be an elegant proof that the llms have outsampled us decisively. one word at a time, which is how we write too.
- bahmboo 2mo agoMicrosoft uses formal verification of their encryption code in production using SymCrypt.
- smalltorch 2mo agoBreaking modern encryption comes down to being in control of key generation rather than brute force. Other than that you'll have a hard time bute forcing 2^256 possibilities. Comes down to a gut feeling but I lean that this stuff is already all figured out.
- tptacek 2mo agoThis is JP Aumasson, the co-author of BLAKE2 and BLAKE3. Aumasson is notorious in cryptography circles for his "too much crypto" argument, that modern symmetric cryptography is overly conservative, running more rounds than are necessary given the very low likelihood that advances in computer science are going make a real dent in them. A distinction a lot of comments in this thread aren't picking up on is the mechanisms that make most asymmetric cryptography work, versus those of symmetric cryptography. Asymmetric constructions like RSA and ECDH are simple mathematical objects, and their security depends on assumptions we make about advanced algebra, number theory, &c. It's plausible to imagine we could discover something about discrete logs that would destabilize DH. It's less plausible to imagine something like that happen to AES, which is deliberately designed not to have clean structure.
- SKYNET800 2mo ago[flagged]
- cootsnuck 2mo agoPlease consider taking a break from your use of LLMs. You are clearly deep in the throes of AI psychosis and need to talk to people you trust in your life instead of the chatbots.
- SKYNET800 2mo agoHave you seen the code and what it does? It’s science, go and take a look.
- ande-mnoc 2mo agoYou are expecting people to read 25K+ lines of code in a single Python file that is generated by LLM and then translate all the comments written in Russian?
- inigyou 2mo agoI did. Looks like technical analysis, which is pseudoscience. For some reason the code also talks about animals and limbs. Very large amounts of the code are also spent on useless details like logging, and monkeypatching matplotlib, that no human would spend so much code on.
- 2mo ago
- biosboiii 2mo agowhy break RSA or AES when you can just subpoena/hack Cloudflare?
- teravor 2mo agosymmetric crypto typically wants the eat the cake and have it too, it wants to be both secure and efficient. to that end, a so-called "security margin" is guessed at and the number of rounds of the cipher is determined accordingly. it is certainly possible for an LLM to prove that the guess was wrong and everything that it implies. having said that, the security of symmetric cryptography relies on the fact that you cannot unwind (find initial conditions) a sufficiently chaotic system in the discrete domain. for example, SHA256 with 512 rounds will almost certainly count as sufficiently chaotic by any definition but it wouldn't be as efficient as the current 64 rounds. it is often said that it's difficult to come up with a secure symmetric cipher on your own, but assuming you know what you are doing it's quite easy. the hard part is to have enough confidence in it to make it efficient.
- xtajv 2mo agoThis is a friendly reminder that although the RSA Factoring Challenge ended in 2007[0], there is still a cool million resting on each of the Millenium Prize Problems. LLMs have not changed the calculus there. [0] https://en.wikipedia.org/wiki/RSA_Factoring_Challenge https://en.wikipedia.org/wiki/RSA_Factoring_Challenge