17 ms·
Linus's transition plan seems to involve truncating SHA-256 to 160-bits. This is bad for several reasons: - Truncating to 160-bits still has a birthday bound a
by bascule 10y ago
Linus's transition plan seems to involve truncating SHA-256 to 160-bits. This is bad for several reasons:
- Truncating to 160-bits still has a birthday bound at 80-bits. That would still require a lot more brute force than the 2^63 computations involved to find this collision, but it is much weaker than is generally considered secure
- Post-quantum, this means there will only be 80-bits of preimage resistance
(Also: if he's going to truncate a hash, he use SHA-512, which will be faster on 64-bit platforms)
Do either of these weak security levels impact Git?
Preimage resistance does matter if we're worried about attackers reversing commit hashes back into their contents. Linus doesn't seem to care about this one, but I think he should.
Collision resistance absolutely matters for the commit signing case, and once again Linus is downplaying this. He starts off talking about how they're not doing that, then halfway through adding a "oh wait but some people do that", then trying to downplay it again by talking about how an attacker would need to influence the original commit.
Of course, this happens all the time: it's called a pull request. Linus insists that prior proper source code review will prevent an attacker who sends you a malicious pull request from being able to pull off a chosen prefix collision. I have doubts about that, especially in any repos containing binary blobs (and especially if those binary blobs are executables)
Linus just doesn't take this stuff seriously. I really wish he would, though.
- tytso 10y agoThat's not the plan. That was an idea that was thrown out if this was an emergency (it's handling different length hashes, and doing so that we don't have to force a flag day conversion which is hard), but once people realized that in fact, the sky was not following, the plan which Linus outlined in his G+ post was devised --- which does not involve truncating a 256-bit hash.
- bascule 10y agoCan you please link me to "the plan" then? I have been trying to follow some of the ML discussion and that was the last plan I saw him put forth, e.g.: https://marc.info/?l=git&m=148787047422954 https://marc.info/?l=git&m=148787047422954
- pedrocr 10y agoBy clicking "next in thread" in that very post you get Linus replying to himself with a non-truncating plan: https://marc.info/?l=git&m=148787163023435&w=2 https://marc.info/?l=git&m=148787163023435&w=2
- bascule 10y agoThanks for the link. Re: snark, I did hit "next in thread" but managed to skim over that in his response. But again, thanks anyway. (Note: this still sounds more like spitballing than "the plan", but at least it's a step in the right direction)
- greglindahl 10y agoDo you have any insight into the sha-256 vs sha-512 choice?
- eslaught 10y agoCan you link the new plan? I didn't see it described in his G+ post.
- yawaramin 10y agoThe G+ post was not meant to describe the plan, it was an overview of the situation for git users who aren't necessarily security experts. The plan is outlined at https://marc.info/?l=git&m=148787163023435&w=2 https://marc.info/?l=git&m=148787163023435&w=2
- deleted 10y ago[deleted]
- ksherlock 10y agoOn the mailing list he suggested displaying 160 bits for backwards compatibility with existing tools. The full 256 bits would be used internally (and displayed externally with the proper --flag)
- diminish 10y agoadditionally cosmetics but instead of hex string could they not use a-zA-Z0-9 in the visual output to the user to make the git command line text output shorter? 0efaa. -> 2AdC..
- pshc 10y agoUnfortunately there's a lot of code out there that assumes commit hashes are case-insensitive and hexadecimal... I've written some myself. It'd be a hugely painful change.
- runeks 10y agoThen you have stuff like 1, I and l, which are difficult to distinguish. Which was why base58 was invented (basically the range you suggested, without visually similar characters).
- kzrdude 10y agoThis really should not be a problem since that's task #1 that a programmer's good font needs to solve.
- cbr 10y agoI often run into git hashes places where a programming font wouldn't apply, like URLs and web pages.
- wongarsu 10y ago>where a programming font wouldn't apply, like URLs Any phisher's wet dream is a URL displayed in a font where l and I look the same. For anything that routinely displays URLs, font matters a lot. And any website displaying git hashes is likely to be programming related and have some code-font (good monospace or similar) in it's repertoir.
- paulddraper 10y ago> especially in any repos containing binary blobs Yeah, that part is the real flaw in the argument.
- hermitdev 10y agoAlthough, I generally defer to someone such as Linus in having far more domain knowledge such as this, but I'm concerned willingness to just drop bits of the hash like this. I mean 20-30 years ago, I could quasi understand for the sake of performance, but are we really so concerned about performance of a few clock cycles vs opening ourselves to a known vulnerable attack?
- hdhzy 10y agoI think the truncation is considered because a lot of software assumes git commit id has 160 bit length. It's not because of performance.
- hermitdev 10y agoThere are only 2 reasonable rationales for the 160 bit restriction: it was arbitrary because that was the length of the chosen hash, or for performance (and no one will ever need more that 640kb of RAM, or 160 bits of hash). I'm guessing it was probably the first, though, and Linus designed for the immediate now, and not for the future.
- derefr 10y agoIt's not even a "restriction", so much as a fragile ecosystem of tools that parse the output of git commands. A lot of them assume particular column widths.
- hueving 10y agoWhat is this known vulnerable attack you are talking about? The birthday paradox applies to every hashing function.
- hermitdev 10y agoWhile there may not be a known, direct, attack against git, WebKit's SVN repository was demonstrably attacked by the SHA-1 vulnerability, oddly by their own developers today (or yesterday?). That it took that little time from a PoC to an actual production issue leads me to believe it wouldn't take long for a dedicated individual or team to extend the vulnerability to git and it's "mitigation" involving the header. I prefer not to take the bury my head in the sand approach to this, at least when it comes to public repositories.
- deleted 10y ago[deleted]
- deleted 10y ago[deleted]
- the_mitsuhiko 10y ago> Linus's transition plan seems to involve truncating SHA-256 to 160-bits. This is not the plan. This is the backwards compatible system for tools that can only deal with 40 characters.
- chainsaw10 10y ago> reversing commit hashes back into their contents Somewhat off topic, but is this actually possible? Given hashing is inherently lossy, I'm inclined to assume it's not possible for anything must longer than a password, but commits are text, which I suppose is low entropy per character, so I don't know.
- bascule 10y agoYes this is possible. It's called a preimage attack: https://en.wikipedia.org/wiki/Preimage_attack https://en.wikipedia.org/wiki/Preimage_attack It's computationally infeasible against most cryptographically secure hash functions. However, Grover's algorithm results in a sqrt(keyspace) reduction in security levels (effectively halving bit-sized security levels): https://en.wikipedia.org/wiki/Grover's_algorithm https://en.wikipedia.org/wiki/Grover's_algorithm
- AgentME 10y agoA successful preimage attack against a cryptographic hash would most likely find you a different message that gets you the same hash. You wouldn't do a preimage attack to find what the original input to create a hash was. You would do a preimage attack to find a new input with the same hash so that you could pass this new input around claiming it was the original (and have any signatures for the old input be valid for your new input, etc).
- oleganza 10y agoWhat you described is called "second preimage attack", that's != "first preimage attack".
- AgentME 10y agoIn a second preimage attack, you start with an input that comes out to a hash, and you're tasked with making a distinct input that comes out to a hash. In a first preimage attack, you start with only the hash, and have to come up with an input from scratch that comes out to that hash. Even in a first preimage attack, you're extremely unlikely to craft the same input that someone else used to create the hash you were given. By pigeonhole principle, the number of inputs that come out to a shorter hash is extreme.
- deleted 10y ago[deleted]
- hueving 10y ago>Preimage resistance does matter Is there a preimage weakness though? I thought this attack only reduced collision resistance.
- OJFord 10y agoGP said that in reference to quantum attacks.
- mappu 10y ago>(Also: if he's going to truncate a hash, he use SHA-512, which will be faster on 64-bit platforms) BLAKE2 is faster still! It's also at least as secure as SHA-3, and produces any choice of output size up to 512bit.
- stouset 10y agoI'm a fan of blake2, but SHA-2 is the conservative choice here given the sheer amount of cryptanalysis it's been subjected to.
- gkya 10y agoIf a repo contains binary blobs, especially executables, well, that's very bad practices right there. Also, how can somebody else modify a binary in a meaningful way and send a patch to it? How can you review a patch to a binary file before applying? I'd say that any sane project, especially if open source, would not include binaries (maybe apart from images), and even if it did, would not accept patches to them (if the members of a project were talking about changing, say, an icon, valid images would be exchanged on a mailing list/issue tracker, nobody would bother making, sending and applying binary diffs; then somebody w/ commit bit just commit it). Let me put it more simply: if you're accepting patches for binary files in your repos you don't care about security at all. Maybe unless if you know how to decode machine code/JPEG manually. Also, the proposition was to use the full hash internally, and truncate the representation. Nobody other than git itself uses full hashes anyways.
- wongarsu 10y agoExecutables in git are bad practise, but not all that uncommon. Images in git are the norm, and if somebody comes in and creates a pull request with an improved version of the existing images (better compression, better adapted for color blind people, fixing whitespace issues etc.) that's pretty unsuspicious and likely to succeed (and I've seen it multiple times).
- usrusr 10y agoBut how would a safer hash improve resistance to that blob update scenario? How would a weaker hash make it more dangerous? If you have a fiercely audited branch next to one that is basically free for all, maybe?
- Dylan16807 10y ago- Truncating to 160-bits still has a birthday bound at 80-bits. That would still require a lot more brute force than the 2^63 computations involved to find this collision, but it is much weaker than is generally considered secure What level is considered secure, then? The numbers for O(time) and O(space) should be many orders of magnitude apart to represent the relative costs.
- simplehuman 10y ago> Linus just doesn't take this stuff seriously. I really wish he would, though. Can't downvote this enough. This is plain FUD. Did you even read the complete thread on the git mailing list? This was just one proposal by him.
- AgentME 10y agoIf he took this stuff seriously, he wouldn't have waited 12 years since SHA-1 was broken to even start considering any changes.
- simplehuman 10y agoSo did anyone on this thread send patched to the mailing list? Obviously you guys are very serious right?
- aseipp 10y agoThis conversation has gone on several times over the course of years on the Git mailing list. In almost every situation it's been completely brushed off as a mostly non-issue to change from SHA-1 inside Git, despite the fact it's been known SHA-1 has basically been on life support. There are lots of opinions on both sides, but ultimately, until now, the Upstream decision seemed to be "WONTFIX". Given this context, of course nobody wrote patches: they would have obviously been rejected and been a total waste of time. Until now, when we actually have to deal with it. This is all aside from your argument being fundamentally weak, however ("you can't criticize anything unless i say so and contributed by meeting this arbitrary standards. i mean, didn't do anything either, i just get to make up the rules you abide by!!")
- simplehuman 10y ago> This is all aside from your argument being fundamentally weak, however ("you can't criticize anything unless i say so and contributed by meeting this arbitrary standards. i mean, didn't do anything either, i just get to make up the rules you abide by!!") Are we both reading the same GP comment? It reads as "If he took this stuff seriously, he wouldn't have waited 12 years since SHA-1 was broken to even start considering any changes.".
- infinity0 10y ago> "A hash that is used for security is basically a statement of trust [..] In contrast, in a project like git, the hash isn't used for "trust". I don't pull on peoples trees because they have a hash of a4d442663580. Our trust is in people, and then we end up having lots of technology measures in place to secure the actual data." This is horseshit, and Linus should not be saying these hugely misleading statements about security principles. The point of a hash is to remove the need for trust between the trusted person who tells you the hash and the infrastructure you get the actual data that was hashed, from (edit: and, between you and the latter). In other words, once you get a good non-colliding hash from a trusted person, then you don't need to worry about malicious infrastructure sending you bad data claiming to be the source of that hash. Linus trusting Tytso to sign the commit object that references the SHA-1 of the tree object, says nothing about whether the infrastructure served him the tree object correctly. Sure, he might also trust the infrastructure providers, but when he says "trusts people" it does not sound like that is what he means. And even if he trusts the infrastructure providers, with a good hash HE DOESN'T HAVE TO. The "trust" wording is serious horseshit. (edit: there is also the case of people downloading "linux" from random git repos in the future. Right now if you GPG-sign a commit or tag, it has SHA-1 references to the tree object underneath it. Once SHA-1 is more broken it basically means you shouldn't trust random git repos across the internet to give you good content, even if it's "signed by Linus".)
- bascule 10y agoYes, exactly. If Linus truly doesn't care about security, then git could use any error correcting code that produces a uniform distribution of tags, such as CRC64. The size of the tag only affects the number of objects we'd expect to be able to commit before we see a collision: over 4 billion in the case of CRC64. Linux mistakenly claims using a cryptographic hash function helps avoid non-malicious collisions, but this is not the case. Where the choice of a cryptographic hash function matters is specifically if we expect an attacker to be trying to collide tags. CRC64 is a linear function of the input data and therefore fails miserably at preventing attackers from colliding tags, but still produces a uniform distribution of tags for non-malicious inputs. git seems to be in the odd place where Linus argues he's using a cryptographic hash function but not for security purposes.
- deleted 10y ago[deleted]
- amluto 10y ago> - Post-quantum, this means there will only be 80-bits of preimage resistance Post-quantum, it's only ~2^53 work to find a collision. IMO that's worse.