3 ms·
To clarify: if you start from a given nonce and key: cipher = AES.new(key, AES.MODE_GCM, nonce) while True: block = f.read(16*1024*1024)
by josephernest 6y ago
To clarify: if you start from a given nonce and key:
cipher = AES.new(key, AES.MODE_GCM, nonce)
while True:
block = f.read(16*1024*1024)
if not block:
break
out.write(cipher.encrypt(block))
you get exactly the same result as if you do (with a big RAM, bigger than your file) it in one pass:
cipher = AES.new(key, AES.MODE_GCM, nonce)
out.write(cipher.encrypt(f.read()))
Please try it with pycryptodome, you will see it is.
You might find this unforunate in the naming, and .init(), .update(), etc. might have been better names to emphasize this.
So this shows that, in its current state, the chunking is just a "RAM-efficient" way to encrypt, but it writes exactly the same encrypted content, as if you did encrypt(...) in one pass. So as long as the file is under ~2^39 bits, it is fine (see https://csrc.nist.gov/publications/detail/sp/800-38d/final https://csrc.nist.gov/publications/detail/sp/800-38d/final).
___
Then, there is another layer of chunking that would be possible, and that would add many benefits: even better deduplication, avoid to reencrypt a whole 10GB file if only a few bytes have changed.
This would be an interesting addition, it's on the Todo list, and I know some other programs do it, of course.
But to clarify: this "content-chunking" is independent to the "RAM-efficient" one I use here.
If you want to continue the discussion (more convenient than here), you're very welcome to post a Github issue.
Thanks for your remarks, I appreciate it.