4 ms·
except ValueError: print('Incorrect key or file corrupted.') So if there is a tag problem, this is clearly logged, and you know something is wrong. The
by josephernest 6y ago
except ValueError:
print('Incorrect key or file corrupted.')
So if there is a tag problem, this is clearly logged, and you know something is wrong. The good thing is that you can easily edit it and have instead (pass the filename fn to the function, to be able to log):
print('Incorrect key or file corrupted, will be deleted:', fn)
os.remove(fn)
exit(...)
Real question: is it possible to `.verify(tag)` before having decrypt(...) the whole file? I doubt it is possible. So an option could be to write the file in a temporary place, and then, only when tag is verified, move it to the right place. Delete it if it is not verified. Another option would be to do a first pass of decrypt(), without writing anything to disk, then get the tag, verify it, and then if ok, redo the whole decryption with writing on disk this time. The latter might be a bit too extreme and halfs the performance.
- rakoo 6y ago> So if there is a tag problem, this is clearly logged, and you know something is wrong [...] The good thing is that you can easily edit it and have instead (pass the filename fn to the function, to be able to log) That's the thing: the script in its current version is incorrect, and even doing that won't be a perfect solution. That's why other people are saying that other softwares, with large usage and that can do more than what nFreezer can do, should be analyzed before trying to do it your own way. It's good to not rely on anyone else, but crypto is the one domain where you can't have "good enough" -- it's either correct, or it's not. > Another option would be to do a first pass of decrypt(), without writing anything to disk, then get the tag, verify it, and then if ok, redo the whole decryption with writing on disk this time Yep, that's the way: do the decrypting in memory, or in /tmp, verify the tag, and only after you can put the file where it belongs. I just checked the API of the crypto module, and there's a `decrypt_and_verify` that should do it properly. Of course that's problematic especially for big files, so what you want to do is chunk the files, encrypt the chunks separately and store the file as a list of such chunks. The step after is to use Content-Defined Chunking, ie chunking based on the content of the file. This way when a big file modifies only the chunk around the modification will change, the rest of the file will be chucked exactly the same way. So you don't need to store the full content of each version of the file, just a small-ish diff. That's not a novel system, bup (https://github.com/bup/bup https://github.com/bup/bup) kinda pioneered it... and as others have advised, restic, borg-backup and tarsnap do exactly that.
- josephernest 6y agoTo clarify: if you start from a given nonce and key: cipher = AES.new(key, AES.MODE_GCM, nonce) while True: block = f.read(16*1024*1024) if not block: break out.write(cipher.encrypt(block)) you get exactly the same result as if you do (with a big RAM, bigger than your file) it in one pass: cipher = AES.new(key, AES.MODE_GCM, nonce) out.write(cipher.encrypt(f.read())) Please try it with pycryptodome, you will see it is. You might find this unforunate in the naming, and .init(), .update(), etc. might have been better names to emphasize this. So this shows that, in its current state, the chunking is just a "RAM-efficient" way to encrypt, but it writes exactly the same encrypted content, as if you did encrypt(...) in one pass. So as long as the file is under ~2^39 bits, it is fine (see https://csrc.nist.gov/publications/detail/sp/800-38d/final https://csrc.nist.gov/publications/detail/sp/800-38d/final). ___ Then, there is another layer of chunking that would be possible, and that would add many benefits: even better deduplication, avoid to reencrypt a whole 10GB file if only a few bytes have changed. This would be an interesting addition, it's on the Todo list, and I know some other programs do it, of course. But to clarify: this "content-chunking" is independent to the "RAM-efficient" one I use here. If you want to continue the discussion (more convenient than here), you're very welcome to post a Github issue. Thanks for your remarks, I appreciate it.
- cperciva 6y agobup (https://github.com/bup/bup https://github.com/bup/bup) kinda pioneered it... and as others have advised, restic, borg-backup and tarsnap do exactly that. According to wikipedia, bup was released in 2010, 3 years after Tarsnap started doing this. (And Tarsnap wasn't the first either.)