10 ms·
Reminder that GitHub has blocked Git commit collisions since 2017, and as far as anybody is aware hasn't seen one in the wild. https://github.blog/2017-03-20-s
by ImJasonH 5y ago
Reminder that GitHub has blocked Git commit collisions since 2017, and as far as anybody is aware hasn't seen one in the wild.
https://github.blog/2017-03-20-sha-1-collision-detection-on-github-com/ https://github.blog/2017-03-20-sha-1-collision-detection-on-...
- api 5y agoRandom collisions in 160-bit space are incredibly unlikely. This is talking about intentional collision, and means that it's entirely feasible for someone with significant compute power to create a git commit that has the exact same hash as another git commit. This could allow someone to silently modify a git commit history to e.g. inject malware or a known "bug" into a piece of software. The modified repository would be indistinguishable if you're only using git hashes. Git's uses SHA-1 for unique identifiers, which is technically okay as long as they are not considered secure. If git were designed today it would probably use SHA2 or SHA3 but it's probably not going to change due to the massive install base. Edit: anyone know if git's PGP signing feature creates a larger hash of the data in the repo? If not maybe git should add a feature where signing is done after the computation of a larger hash such as SHA-512 over all commits since the previous signature.
- asdfman123 5y ago> Random collisions in 160-bit space are incredibly unlikely Just for fun: to get a 5% chance of a hash collision between ANY two numbers in an 160 bit space, you'd have to generate 3.9e23 hashes. So you'd have to generate 1000 hashes per second for *12 trillion years.* Formula: n = sqrt(2 * 2^160 * ln(1/(1-0.05)) https://en.wikipedia.org/wiki/Birthday_problem#Probability_of_a_shared_birthday_(collision) https://en.wikipedia.org/wiki/Birthday_problem#Probability_o...
- throw_away 5y agoLooks like there's a migration plan for git to SHA-256: https://git-scm.com/docs/hash-function-transition/ https://git-scm.com/docs/hash-function-transition/
- lokedhs 5y agoLast updated in 2017. What is the current status of this?
- beermonster 5y agoHere is discussion [1] on the issue on the Git mailing list at the time with some useful context WRT git. Unsure what current status is. Git 2.29 Had experimental support for sha256.. not sure what it’s current status is, but that was a year ago. [1] https://marc.info/?t=148786884600001&r=1&w=2 https://marc.info/?t=148786884600001&r=1&w=2
- Greek0 5y agoAs of January 2021, SHA256 repositories are supported, but experimental. They can be created with `git init --object-format sha256`. If I understand correctly, they don't mix at all with SHA1 repositories (i.e. you can't pull/push from/to between SHA1 and SHA256 repos). See https://stackoverflow.com/questions/65870508/git-and-sha-256 https://stackoverflow.com/questions/65870508/git-and-sha-256
- a_t48 5y agoAs far as I know it only signs individual commits.
- ImJasonH 5y agoYep. GitHub is saying they would block an object that looked like it was crafted to produce a collision using SHAttered, and hasn't seen it, intentionally or otherwise, in the wild.
- deleted 5y ago[deleted]
- AnimalMuppet 5y agoCan you create an intentional collision, and also have the counterfeit have compilable code? (Or, in other circumstances, intelligible English text?)
- jfrunyon 5y agoYes, just like you can create an intentional collision, and also have the counterfeit be a PDF. [https://shattered.io/ https://shattered.io/]
- deleted 5y ago[deleted]
- jfrunyon 5y agohttps://stackoverflow.com/questions/23584990/what-data-is-being-signed-when-you-git-commit-gpg-sign-key-id https://stackoverflow.com/questions/23584990/what-data-is-be... says git signing just signs the commit hash (plus metadata: committer, commit message, etc).
- oconnore 5y agoI was surprise that no one suggested truncating SHA-256 to 160 bits (same as for SHA2-256/224, or SHA2-512/256). The attacks on SHA-1 are not directly based on the length of the hash, they are based on weaknesses in the algorithm. Even attacking SHA2-256/128 would be quite difficult as I understand it, even though it's the same length as MD5. Truncated hashes also of course have the great property that they mitigate the length extension in Merkle-Damgard
- duskwuff 5y ago> I was surprise that no one suggested truncating SHA-256 to 160 bits... To do that, you have to stop generating commit hashes with SHA-1, breaking compatibility with existing git clients. And if you're going to do that, you might as well just use the whole SHA-256 hash, since compatibility is out the window already.
- a1369209993 5y ago> Truncated hashes also of course have the great property that they mitigate the length extension in Merkle-Damgard To be fair, this is totally irrelevant to git, since the attacker knows the whole message and can just recompute the extra bits themselves. That said: > I was surprise that no one suggested truncating SHA-256 to 160 bits (same as for SHA2-256/224, or SHA2-512/256). The attacks on SHA-1 are not directly based on the length of the hash, they are based on weaknesses in the algorithm. Very seconded. You could even shove the extra 96 bits in a optional metadata field and have new versions of git throw up a giant air-raid-siren-level error if they don't match (since that will never happen by accident[0]) and still have the full 256-bit-hash worth of security for most purposes. Git already allows (arguably encourages) people to truncate hashes to 28 bits or so at the UI level, so there's precendent for that already. 0: You do not have anywhere near 2^80 commits in the world, much less in the same repo.
- adrian_b 5y agoIt is not yet possible "to create a git commit that has the exact same hash as another git commit" in the sense that if someone else has already done a commit you can make another commit with the same hash. What is possible now is something that is much easier: if you have enough money and time, you can create 2 commits with the same hash, which start with some different parts, which may be chosen arbitrarily, then they have to include some parts that must be computed to ensure that the hashes will match and which will be gibberish that might be disguised in some constants or some binary blob, if possible. Then you can commit one of them and presumably you can later substitute the initial commit with the other one without anybody being able to detect the substitution.
- outworlder 5y ago> presumably you can later substitute the initial commit with the other one without anybody being able to detect the substitution. How? Which operation would be involved? Will it not show anywhere else(reflog)?
- slim 5y agoBy replacing the file on disk (you need to have the capability of acquiring access to the server)
- adrian_b 5y agoBecause such an operation is unlikely to succeed even if the hashes match, replacing SHA-1 was not a high priority for Git.
- noway421 5y ago`git push --force` I assume. In GitHub, this will leave a trail of events though - Webhook events and auditing.
- cerved 5y agoforce push only updates a remote ref if it's not a fast forward. it doesn't do any forced pushing of objects. you need to modify the object store manually, so access to the filesystem
- outworlder 5y ago> This could allow someone to silently modify a git commit history to e.g. inject malware or a known "bug" into a piece of software. You need a collision. You also need it to be syntactically correct. You need it to not raise any red flags if you are contributing a patch. And ultimately you need it to do what you want. That's a pretty tall order.
- api 5y agoAll you need is one variable sufficiently large that you can change until you find a collision, such as a comment (which about all languages allow) or a meaningless number. You could even vary whitespace until it fits, like spaces at the end of lines.
- dude187 5y agoYou'd also need the actual patch to survive future commits, especially without introducing any merge conflicts
- sirclueless 5y agoCommits aren't patches. They contain the whole tree. Retroactively changing a commit can't possibly introduce conflicts with other commits on top of it, the worst it can do is introduce big funny-looking diffs.
- dude187 5y agoWell, the contents of the commit is a patch plus metadata. They point to a parent commit, and layer themselves in the tree. The problem would be if a clone doesn't fetch the new version of the patch and generates a new commit that would conflict with the modified commit. You're changing the base all the future diffs are based off of. It might just jumble the source essentially corrupt the file, but I'm not sure.
- wheybags 5y agoContents of a commit is not a patch, it's the whole tree. The got ui presents it as a patch, but that's generated at runtime by the "git diff" command. It does internally use delta compression to save storage space, but it's not necessarily a straight delta between a commit and its direct ancestor (and that's just an internal optimisation).
- noway421 5y agoFrom a practical point of view, how would injecting malware happen? If you're trying to insert a malicious diff somewhere deep in the git history, you would need to recompute all the other commits after the injected commit - which would most certainly change their commit ids too if they are touching the same file. When other commit ids change, the malicious change becomes detectable. There's also the case for auditing: force pushing into an existing repo triggers an event in GitHub and is logged. While the logging event can be missed, it leaves a paper trail. With things like reproduce-able builds, this also becomes harder. Distributing (through a means of a fork, or putting it up on a website mytotallylegitgitrepos.com) source code which builds into a binary which doesn't match upstream hash is suspicious.
- ufo 5y agoPresumably the attacker would modify the most recent commit which edits the file that they are targeting. It is true that the attack becomes more difficult if you try to target an older commit. Auditing helps if they try to force push the original repo, but doesn't protect vs someone redistributing malicious clones of the repo. Reproduceable builds do help, but only for projects that can take advantage of it...
- sirclueless 5y agoEvery commit references every file. If you change the content of an old commit you would only affect people who check out the old commit. So this is utterly pointless and not what someone would do. Instead what you would do is attempt to make a file-object that has a certain SHA1 hash identifying it, and a colliding file-object that has the same SHA1 hash. Then you are free to give people who clone the repository different file contents depending on when/who/how someone requests it (if the file content is hosted on github, how to change the file object identified by a given SHA1 hash is an additional hurdle since it's assumed to be immutable and indefinitely cacheable; if you control the host yourself you can just change it whenever you like).
- tialaramex 5y agoThe defence used by GitHub specifically defends against these intentional collisions, not some mirage of random collisions. Basically you collide a hash like SHA-1 or MD5 by getting it into a state where transitions don't twiddle as many bits, and then smashing the remaining bits by brute force trial. But, such states are weird so from inside the hash algorithm you can notice "Huh, this is that weird state I care about" and flag that at a cost of making the algorithm a little slower. The tweaked SHA1 code is publicly available. If you're thinking "Oh! I should rip out our safe SHA256 code and use this unsafe but then retro-actively safer SHA1" No. Don't do that. SHA-256 is safer and faster. This is an emergency patch for people for whom apparently 20 years notice wasn't enough warning. In theory the known way to do this isn't the only way, but, we have re-assuring evidence for MD5 that independent forces (probably the NSA) who have every reason to choose a different way to attack the hash to avoid detection do trigger the same weird states even though they're spending the eye-watering sum of money to break hashes themselves not just copy-pasting a result from a published paper.
- api 5y agoGit != GitHub
- RicoElectrico 5y agoSo, if I understand correctly: the patched SHA-1 code generates the same hash, but is has checks on the internal state so that it will flag inputs which are likely to be intentionally colliding?
- versteegen 5y agoYes. https://github.com/cr-marcstevens/sha1collisiondetection https://github.com/cr-marcstevens/sha1collisiondetection
- stevekemp 5y ago> If git were designed today it would probably use SHA2 or SHA3 but it's probably not going to change due to the massive install base. There is some work going on to change this, but it's not an easy task: https://lwn.net/Articles/811068/ https://lwn.net/Articles/811068/
- tantalor 5y ago> A higher probability exists that every member of your programming team will be attacked and killed by wolves in unrelated incidents on the same night. - Scott Chacon
- asdfman123 5y agoThat's even true if you make your developers commute into Yellowstone National Park by ski every day
- bawolff 5y agoThis is stupid. The probability of any attack happening at random is close to zero. E.g. the probability of a buffer overflowing just right to give you a shell through random chance is less than being eaten by wolves. The question is about the risk of someone intentionally performing the attack, not the probability it will accidentally happen at random.
- kragen 5y agoThis turns out to be wrong; for a 6-member programming team, that probability is about 2⁻²⁴⁵, which is about 2⁸⁵·³ times less likely than an accidental 160-bit SHA-1 collision: http://canonical.org/~kragen/sw/dev3/rpn-edit#3_8_0_1_0_0_0_7_0_0_+_+_+_+_+_+_+_+_+_2015_2000_-_365_1_4_/_+_*_7_1000_1000_1000_*_*_*_*_/_ln_6_*_2_ln_/ http://canonical.org/~kragen/sw/dev3/rpn-edit#3_8_0_1_0_0_0_... Aside from being bullshit, it's also irrelevant, since we're discussing a collision being generated on purpose, not by accident.
- CobrastanJorji 5y agoJust to nitpick, I don't think that formula is valid. We're primarily interested in "unrelated" wolf attacks, but it counts the total fatalities, not the total number of fatal incidents. If we count each fatal attack as only one incident, regardless of the casualties, we get 2^-258 instead. But of course we also need to take into account where the 6-member team lives. If they all live in West Bengal, India, the consideration is much different than if our developers live in Atlanta. Atlanta doesn't have any wild wolves. There is a Wolf's guenon in the zoo, but that probably doesn't count as a risk because they mostly eat small animals and also are monkeys.