5 ms·
Just to be clear: Every time you modify a file, the new changes get put in using SHA3. In an older repository, any given commit might have some files identifie
by SQLite 7y ago
Just to be clear:
Every time you modify a file, the new changes get put in using
SHA3. In an older repository, any given commit might have some
files identified using SHA1 (assuming they have not changed
in 3 years) and others identified using SHA3.
For example, the manifest of the latest SQLite check-in is
see at (https://www.sqlite.org/src/artifact/29a969d6b1709b80 https://www.sqlite.org/src/artifact/29a969d6b1709b80).
You can see that most of the files have longer SHA3 hashes,
but some of the files that have not been touched in three
years still carry SHA1 hashes.
An attack like what you describe is possible if you could
generate an evil.c file that has the exact same SHA1 hash as the
older floppy.c file. Then you could substitute the evil.c
artifact in place of the floppy.c artifact, get some
unsuspecting victim to clone your modified repository, and
cause mischief that way. Note, however, that this is a
pre-image attack, which is rather more difficult to pull off
than the collision attacks against SHA1, and (to my knowledge)
has never been publicly demonstrated. Furthermore, the evil.c
file with the same SHA1 hash would need to be valid C code that
does something evil while still yielding the same hash (good
luck with that!) and Fossil (like Git) has also switched over
to Hardened SHA1, making the attack even harder still.
As still more defense, Fossil also maintains a MD5 hash against
the entire content of the commit. So, in addition to finding
evil.c that compiles, does your evil bidding, has the same
hardened-SHA1 hash as floppy.c, you also have to make sure that
the entire commit has the same MD5 hash after substituting
the text of evil.c in place of floppy.c.
So, no, it is not really practical to hack a Fossil repository
as you describe.
- wyoung2 7y ago> Furthermore, the evil.c file with the same SHA1 hash would need to be valid C code that does something evil while still yielding the same hash ...and also produce an innocent-looking diff! I mean, you could stuff a bunch of random bytes into a C comment to force the desired hash in the output using these documented attack techniques, but anyone inspecting the diffs between versions is likely to see such an explosion of noise and call foul. If you want an analogy, it's like someone saying they've learned to impersonate federal agent identification cards, only it requires that the person carrying the fake ID to have a thousand rainbow-dyed ducks on a leash in tow behind him. Such attacks are fine when it's dumb software systems doing the checks, but for a source code repository where people do in fact visually check the diffs occasionally? Well, let's just say that when someone manages to use SHAttered and/or SHAmbles type attacks on Git (or even Fossil) I expect that it won't take a genius detective to see that the repo's been attacked.
- tjoff 7y agoMany diff tools don't highlight whitespace-only changes. Or at least not in a clear manner. Also, if something is replaced in the history how often do people go back and view diffs in old code? Hardly often enough to rely on it being spotted.
- wyoung2 7y agoIt only takes one person to raise the flag. Sure, many thousands of people doing blind "git clone && configure && sudo make install" could be burned by a problem like this, but someone would eventually do a diff and see the problem on any project big enough to have those thousands of trusting users in the first place. I'm not excusing these SHA-1 weaknesses, only pointing out that it won't be trivial to apply them to program source code repos no matter how cheap the attacks get. For instance, the demonstration case for SHAttered was a pair of PDFs: humans can't reasonably inspect those to find whatever noise had to be stuffed into them to achieve the result. I also understand that these SHA-1 weaknesses have been used to attack X.509 certificates, but there again you have a case very unlike a software code repo, where the one doing the checking isn't another programmer but a program.
- remram 7y agoThe problem is that we are considering an issue where different people can get different objects for the same hash. If the people checking all see the valid files, they cannot raise any alarms to save the poor victims who got poisoned with the wrong objects. They'll clone from the wrong fork, and no amount of checking hashes or signed tags will prevent them from running compromised code.
- wyoung2 7y ago> If the people checking all see the valid files ...which will likely contain thousands of bytes of pseudorandom data in order to force the hash collision... > they cannot raise any alarms You think a human won't be able to notice that the diff from the last version they tested looks awfully funny? Code that can fool the compiler into producing an evil binary is one thing, but code that can pass a human code review is quite another. You might be surprised how often that occurs. I don't do a diff before each third-party DVCS repo pull, but I do diff the code when integrating such third-party code into my projects, if only so I understand what they've done since the last time I updated. Commit messages, ChangeLogs, and release announcements only get you so far. Back when I was producing binary packages for a popular software distribution, I'd often be forced to diff the code when producing new binaries, since several of the popular binary package distribution systems are based on patches atop pristine upstream source packages. (RPM, DEB, Cygwin packages...) Each time a binary package creator updates, there's a good chance they've had to diff the versions to work out how to apply their old distro-specific patches atop the new codebase. Someone's going to notice the first time this happens, and my guess is that it'll happen rather quickly.
- apeace 7y agoAnd if you are that concerned about this type of attack, it may be worth your time to simply start a new Fossil repository using the sha3-only hash policy (writing a script to replay commits into the new repo, so you don't lose history). It seems like a problem very few people need to worry about and Fossil has made the right trade-offs.
- seniorsassycat 7y agoIsn't this the same attack given as an example why git is migrating hash functions in the subject article? The attack may be difficult and unlikely I'm not questioning that, but if I understand correctly then Fossil's migration is straightforward because they did not address the same issues Git chose to.
- SQLite 7y ago> if I understand correctly then Fossil's migration is straightforward because they did not address the same issues Git chose to. I think more is at play here. (1) You can set Fossil to ignore all SHA1 artifacts using the "shun-sha1" hash policy. (2) The excess complication in the Git migration strategy is likely due to the inability of the underlying Git file formats to handle two different hash algorithms in the same repository at the same time. But, I could be wrong. Post a rebuttal if you have evidence to the contrary.
- mb7733 7y ago(2) The excess complication in the Git migration strategy is likely due to the inability of the underlying Git file formats to handle two different hash algorithms in the same repository at the same time. But, I could be wrong. Post a rebuttal if you have evidence to the contrary. It seems unfair to demand a rebuttal when you are the one who made the claim. According to the article at least, the difficulty stems mainly from their migration strategy, for converting all existing SHA1 hashes.
- velcrovan 7y ago> the difficulty stems mainly from their migration strategy, for converting all existing SHA1 hashes. That's essentially the same difficulty, since the only strategy for doing this that has been historically proven to work seamlessly and painlessly involves being able to handle both hash algorithms in the same repository at the same time.
- mb7733 7y agoDoesn't all of this apply to git just as well, except for the last bit about the MD5 hash? It just seems to me that the Fossil maintainers have decided that keeping all old SHA1 hashes is acceptable, while the git maintainers have decided that it is not. Unless I've misunderstood, this is why it was "so easy" for Fossil to transition to a new hashing algorithm. Not some superiority in the design of Fossil, as implied on the Fossil forums.