6 ms·
> "A hash that is used for security is basically a statement of trust [..] In contrast, in a project like git, the hash isn't used for "trust". I don't pull on
by infinity0 10y ago
> "A hash that is used for security is basically a statement of trust [..] In contrast, in a project like git, the hash isn't used for "trust". I don't pull on peoples trees because they have a hash of a4d442663580. Our trust is in people, and then we end up having lots of technology measures in place to secure the actual data."
This is horseshit, and Linus should not be saying these hugely misleading statements about security principles.
The point of a hash is to remove the need for trust between the trusted person who tells you the hash and the infrastructure you get the actual data that was hashed, from (edit: and, between you and the latter).
In other words, once you get a good non-colliding hash from a trusted person, then you don't need to worry about malicious infrastructure sending you bad data claiming to be the source of that hash.
Linus trusting Tytso to sign the commit object that references the SHA-1 of the tree object, says nothing about whether the infrastructure served him the tree object correctly. Sure, he might also trust the infrastructure providers, but when he says "trusts people" it does not sound like that is what he means. And even if he trusts the infrastructure providers, with a good hash HE DOESN'T HAVE TO.
The "trust" wording is serious horseshit.
(edit: there is also the case of people downloading "linux" from random git repos in the future. Right now if you GPG-sign a commit or tag, it has SHA-1 references to the tree object underneath it. Once SHA-1 is more broken it basically means you shouldn't trust random git repos across the internet to give you good content, even if it's "signed by Linus".)
- bascule 10y agoYes, exactly. If Linus truly doesn't care about security, then git could use any error correcting code that produces a uniform distribution of tags, such as CRC64. The size of the tag only affects the number of objects we'd expect to be able to commit before we see a collision: over 4 billion in the case of CRC64. Linux mistakenly claims using a cryptographic hash function helps avoid non-malicious collisions, but this is not the case. Where the choice of a cryptographic hash function matters is specifically if we expect an attacker to be trying to collide tags. CRC64 is a linear function of the input data and therefore fails miserably at preventing attackers from colliding tags, but still produces a uniform distribution of tags for non-malicious inputs. git seems to be in the odd place where Linus argues he's using a cryptographic hash function but not for security purposes.
- Anderkent 10y agoWell, if the cost of computation is not too relevant, and if you don't explicitly need the ability to craft collisions, why would you use a non-cryptographic hash function? Like, when I'm building a lookup index for files, I'm going to use sha-(something), because it's easy and well known. I don't particularly care about the security aspect; I care that everyone immediately knows the contract of sha-1.
- bascule 10y agoThere is nothing to be gained from using cryptographic primitives in a non-security context. You could just as easily use e.g. CRC32 for the case you're describing. There is, however, a performance cost in using cryptographic primitives in non-security-related contexts. You may not care about performance, but it certainly matters for something like git. Linus claims: "So in git, the hash is used for de-duplication and error detection, and the 'cryptographic' nature is mainly because a cryptographic hash is really good at those things." CRC produces a distribution just as uniform as a cryptographic hash function, and it's faster to boot. If these are the only things he actually cares about, and he's explicitly discounting security, he's choosing a slower primitive for no reason. He writes off CRC inexplicably earlier in the post: "Other SCM's have used things like CRC's for error detection, although honestly the most common error handling method in most SCM's tends to be 'tough luck, maybe your data is there, maybe it isn't, I don't care'." Linus seems to think that SHA1 has some sort of magic crypto sauce which magically makes the distribution it produces more uniform than CRC's. It doesn't. The only difference is SHA1 was originally designed to be resistant to preimage and collision attacks, both of which are irrelevant outside of a security context.
- pvg 10y agoYou can still make a not-totally-unreasonable argument that something like CRC64 is simply too small - that maybe the 1 in a million collision chance for a few million hashes is too high. The fast, keyed 'semi-cryptographic' big hashes that are common now weren't around when git was written so the easiest thing to reach for would have been something like SHA-1.
- ploxiln 10y agoYou're saying Linus's statements are "hugely misleading", but it's just that you wish git were designed to be used differently. So, your argument is "horseshit". Linus could have designed a cryptographically perfect system such that he could pull Tytso's signed commit from anywhere on the internet - but he didn't Linus used sha1 as a useful tool for an effective DVCS with an initially simple implementation. He still depends on the security of the kernel.org servers, his work computers, and the top submaintainer's work computers. His git trees and all the submaintainers he pulls from are hosted on kernel.org servers. Security of the kernel.org servers is taken very seriously, especially since the well-known break-in a few years ago. Now even two-factor auth is involved in all git pushes to kernel.org servers. Sub-maintainers make pull-requests by branch name - "please pull branch for-linus ...". More peripheral contributors submit their work via patches on LKML, no commit hashes involved. Finally, no well-known SCM previous to git was based on perfect cryptographic proof of source history, or anything like that. It wasn't a big issue, and it's not the problem git focused on solving. Before git, we all used CVS, SVN, tarballs, and patches. And a significant portion of developers did not use any VCS at all. How could any of us have trusted any source code before 2005?! Somehow we did, though...
- paulddraper 10y ago> Before git, we all used CVS, SVN, tarballs, and patches Yeah, and that sucked. Git is this close (I'm holding my fingers very close together) from providing cryptographic proof of source history. And that's why people think it should be included.
- infinity0 10y agoWhy go through all the effort of securing kernel.org etc etc etc when a simple s/SHA1/SHA512/g on git would suffice? (Ignoring the non-git infrastructure there, for the purposes of this argument.) Not just for him, but for everyone in the world who can't download directly from kernel.org? What's the point of GPG signing tags and commits and adding these as git features, when SHA1 pointers make it pointless? If the cryptographic strength of the hash doesn't matter, why bother even talking about changing the hash from SHA1 to SHA3-256? It's a huge shame because git is very close to being a cryptographically perfect system, but the original creator can't get his security reasoning correct, and does so so badly yet so publicly, and minions who don't know what they're talking about come flocking to defend this.
- azernik 10y agoThis should really be a top-level comment.