6 ms·
Further details as to why Torvalds is not concerned: From the email... "I haven't seen the attack yet, but git doesn't actually just hash the data, it does pr
by groovybits 7y ago
Further details as to why Torvalds is not concerned:
From the email...
"I haven't seen the attack yet, but git doesn't actually just hash the
data, it does prepend a type/length field to it. That usually tends to
make collision attacks much harder, because you either have to make
the resulting size the same too, or you have to be able to also edit
the size field in the header."
[...]
"I haven't seen the attack details, but I bet
(a) the fact that we have a separate size encoding makes it much
harder to do on git objects in the first place
(b) we can probably easily add some extra sanity checks to the opaque
data we do have, to make it much harder to do the hiding of random
data that these attacks pretty much always depend on."
- wyldfire 7y agoI think those are pretty practical approaches. But it sounds as if the cost of changing the hash algorithm is high. What are the impacts of this change? How many things would break if git just changed the algorithm with each new release? Does git assume that the hash algorithm is statically given to be SHA-1 or are there qualifiers on which algorithm is enabled/permitted/configured?
- paulddraper 7y agoAfter making the actual code change, the biggest problem is breaking compatibility with decades of tools in the ecosystem that rely on historically consistent SHA-1 hashes. Git is moving to a flexible hash though. [1] [1] https://stackoverflow.com/questions/28159071/why-doesnt-git-use-more-modern-sha/47838703#47838703 https://stackoverflow.com/questions/28159071/why-doesnt-git-...
- toyg 7y agoMaybe it's time for a version 3 that breaks a bit of compatibility. The Python community would freak out, lol.
- simias 7y agoThe cost is very high but it's only getting higher with time. People have known that SHA-1 was weak and deprecated for much of git's existence. Doing the switch in 2010 would've been painful, doing it now would be orders of magnitude more so and I doubt it'll get any easier in 2030 unless some other SCM manages to overtake git in popularity which seems unlikely at this point. Unless Linus really believes that git will be fine using SHA-1 for decades to come I don't think it's very responsible to keep kicking the ball down the road waiting for the inevitable day when a viable proof of concept attack on git will be published and people will have to emergency-patch everything.
- est31 7y agoThat first quote is misleading. git's special hashing scheme doesn't make the attack "much harder". First there is no difference in length in the original shattered collision already: $ curl https://shattered.io/static/shattered-1.pdf https://shattered.io/static/shattered-1.pdf | wc -c 422435 $ curl -s https://shattered.io/static/shattered-2.pdf https://shattered.io/static/shattered-2.pdf | wc -c 422435 Second, the length is already being hashed into the content during computation of a SHA-1 hash. Look up Merkle-Damgard construction: https://en.wikipedia.org/wiki/Merkle%E2%80%93Damg%C3%A5rd_construction https://en.wikipedia.org/wiki/Merkle%E2%80%93Damg%C3%A5rd_co... There is benefit in storing the length at the prefix as well, as you can avoid length extension attacks, but that's not making attacks "much harder".
- tny 7y agoBut even if the lengths are same, the resulting SHA1 will be different since you prefix the length before hashing
- est31 7y agoThe shattered prefix was chosen as well, see my other comment in the thread: https://news.ycombinator.com/item?id=21980759 https://news.ycombinator.com/item?id=21980759 The only thing that prefixing the length makes difficult is using the same prefix multiple times: you basically have to make up your mind about the type and length before mounting the shattered attack. Also, the prefix means you have to do your own shattered attack and can't use the PDFs that google provided as proof of their project's success. Price tag for that seems to be 11k. [1]: https://github.com/cr-marcstevens/sha1collisiondetection https://github.com/cr-marcstevens/sha1collisiondetection
- toyg 7y agoYeah, that quote doesn't exactly make me confident about Linus's understanding of this particular issue.
- DarkWiiPlayer 7y agoIt's not about it being the same length, but the length of the data being part of the hashed data, which, Linus assumes, will likely make it more difficult to find a collision. He even says at the beginning that he hasn't had a look at the attack yet and is just making an assumption.
- simias 7y agoThe fact that this attack is chosen prefix does weaken the first argument though, you may now find a collision even accounting for any prefixed git "header". The rest is still completely valid though. I still feel like they really should've taken this problem more seriously and earlier. The more we wait the more painful the migration will be when the day comes to move to a different hash function, because everybody knows that'll happen sooner or later. Two years ago we had a collision, now we have chosen prefix, how much longer until somebody actually manages to make a git object collision? And keep in mind that public research is probably several years behind top secret state agency capabilities. Let's stop looking for excuses every time SHA-1 takes a hit and rip the bandaid already. It's going to be messy and painful but it has to be done.
- deleted 7y ago[deleted]
- bjornsing 7y ago> you have to be able to also edit the size field in the header.” As I read the OP [1] a chosen-prefix collision attack such as this allows you to “edit the size field in the header”. Or am I missing something? 1. “A chosen-prefix collision is a more constrained (and much more difficult to obtain) type of collision, where two message prefixes P and P’ are first given as challenge to the adversary, and his goal is then to compute two messages M and M’ such that H(P || M) = H(P’ || M’), where || denotes concatenation.” EDIT: On second thought I was missing something: the adversary is further constrained in the git case because it must find M and M’ of correct length (specified in P and P’). Linus is right (as usual), this probably makes it much harder.
- bjornsing 7y agoA few emails forward in the thread Linus explains though why we don’t need to worry much about this attack in practice: https://marc.info/?l=git&m=148787287624049&w=2 https://marc.info/?l=git&m=148787287624049&w=2 This argument sounds sound to me.
- Thorrez 7y agoHis argument assumes the file is text and people read the entire thing. If either of those assumptions are false, then it's not safe. People store things in git that aren't text. Therefore it's not safe.
- tedunangst 7y agoWhat if somebody makes an attack where they can choose the size and then find a collision?
- banana_giraffe 7y agoLike the two files on the linked page?
- Skunkleton 7y agoThe two files on the linked page were both full of junk data. I suspect that those files being of the same length isn't the norm.
- joeyh 7y ago"(b)" is kind of amusing.. It's been known since 2011 that collision generating garbage material can be put after a NUL in a git commit message and hidden from git log, git show, etc. Still not fixed. With this chosen-prefix attack, they chose two prefixes and generated collisions by appending some data. So your two prefixes just need to be "tree {GOOD,BAD}\nauthor foo\n\nmerge me\0" The only thing preventing injecting a backdoor into a pull request now seems to be git's use of hardened sha1.