13 ms·
The Curious Case of MD5
- kozak 3y agoI still use MD5 as a 128-bit checksum algorithm that is fast and universally supported and compatible everywhere. In this role it's still useful, just don't expect it to be a cryptographic hash anymore.
- danpalmer 3y agoWhile I see the point, what starts as a checksum can easily become relied upon for security over time, after all, checking whether bits have been modified accidentally on purpose, is a subtle distinction in many systems. SHA256 is also near universally supported and doesn’t have this drawback. The only cases where MD5 would be available and SHA256 wouldn’t, is systems that are out of security support anyway, where there are bigger problems to contend with.
- hoten 3y agoSHA256 is something like 30 percent slower than MD5. I'd suggest using Adler (what zlib does) for a simple and fast checksum. Then that should, one hopes, be painfully obvious to be a bad fit for anything security related.
- AlienRobot 3y agoIn Python the Adler library returns a 32 bit checksum. It works pretty well when you're comparing one file to another file. It doesn't work pretty well if you want to, for example, create a quick fingerprint that (tries to) uniquely identify tens of thousands of files. On StackOverflow I saw someone say that they got hash collisions in MD5 (128 bit) after hashing around 20k files. When I tried making something similar I figured if I added the size of the file in bytes to the hash that would decrease the number of hash collisions since you would need a permutation of bytes in a set of bytes of same size to generate the same MD5 hash to get a collision. Still feels random and unavoidable in the greater scheme of things, though.
- magicalhippo 3y agoWhen making a dupe detector some decades ago, I kept the filesize outside of the hash. No need to compute the hash until there's at least two files with the same size.
- gojomo 3y agoConventional analysis of non-contrived files would suggest only a 50% chance of an MD5 collision among 2^64 (18 quintillion) files. So any SO account of "hash collisions in MD5 (128 bit) after hashing around 20k files" may be indicative of some other bug, misreporting, or having a set of files that includes pairs intentionally-contrived to include MD5 collisions. (That's now trivial as long as you're not targeting a specific MD5 value, just intending two files to match, and have some ranges in the files where you can stuff the right result-aligning binary patterns.)
- AlienRobot 3y agoI'm not well-versed in this, but I think you're wrong. A MD5 hash has 2^128 permutations, so it's like the birthday problem, you'll get 50% chance of any two files in a set having the same hash with far fewer than 2^64 files. Besides, we're talking about probabilities. It's possible for 2 files, one after the other, to have the exact same hash just because of chance. In fact, I think there was a bitcoin bug due to this, something about making the wrong assumption that the hash of the entire blockchain and the hash of the entire blockchain + 1 block would never be the same hash, if I remember correctly. Hashes are simply not a reliable method of uniquely identifying things. They are a convenient method but if you implement any hash-based system for identification you should also implement a escape hatch for when two different things have the same hash but must be treated differently.
- gojomo 3y agoNope. 2^64 is the number of tries needed for a 50% chance of a collision between two of the tries. You don't have to stay poorly read on this, and groundlessly allege wrongness: https://auth0.com/blog/amp/birthday-attacks-collisions-and-password-strength/ https://auth0.com/blog/amp/birthday-attacks-collisions-and-p...
- adrian_b 3y agoOn a modern CPU (i.e. 64-bit Arm since 2012, Intel Atom since 2016, AMD Zen since 2017, Intel Core since 2019) SHA-256 is twice faster than MD5. The difference in speed between the hashing speeds is actually greater, a double speed for a long file is what you get when the execution time includes parts that are identical for the two hashes, i.e. launching md5sum/sha256sum and reading the file. Older versions of the binary coreutils package may mask this speed difference by having executables compiled only for very old CPUs. Recent coreutils versions normally use for hashing the OpenSSL library, if found, and OpenSSL uses the hardware instructions where available. Where a bad sha256sum is installed, "openssl dgst -r -sha256" should work instead.
- jeffparsons 3y agoThis happened to me. Users initially couldn't directly control the content being hashed, because it contained a random element (via UUID). Later, the API surface expanded. Luckily, my personal rule is to default to a cryptographic hash unless I can convince myself that cryptographic robustness will never matter and performance definitely will matter, rather than the other way around. In this specific case all users were internal to the company, so it wouldn't have really mattered if it was vulnerable. But it could just have easily been an external user-facing thing.
- ms512 3y agoWhen you don't have a need for a cryptographic digest, it's important to think of the channel's bit error distribution in selecting a checksum algorithm. Different checksum algorithms can provide better error detection for specific channel error models (potentially even with fewer bits). Non-cryptographic checksums are typically designed for various failure models like a burst of corrupt bits, trading off what they do/don't detect to better match detection of corruption in the data they will protect. For example, if you know that there will be at most one bit flip in your message, a single bit checksum (parity check) is sufficient to identify that an error occurred, regardless of your message size. (Note that this is an illustrative example only, since, typically, messages have a certain number of errors for a certain number of message bits -- the expected number of errors depends on the size of the message.)
- drisden84 3y ago"When you don't have a need for a cryptographic digest, it's important to think of the channel's bit error distribution in selecting a checksum algorithm." Important real-life-facts. There was no "give-me-an-appropriate-hash" function. There was: md5sum yourfile.txt Nobody wants to think about "channel's bit error distribution" in a non-security critical context. In fact, its irrelevant, and possibly a usability issue.
- charlieyu1 3y agoYou may as well use CRC32 or Alder32.
- adrian_b 3y agoThose are suitable for error detection for relatively short files, in the kilobyte range. They can be used for error detection in big files only if you compute one per page, e.g. one for each 4kB page. They are not useful as file identifiers. I have found multiple CRC32 collisions even in a single directory (a big one, with around ten thousand files). For error detection in a big file, you need at least some 64-bit CRC, though SipHash is likely to be a better choice than a CRC. For identifying uniquely a file in a multi-TB file system, which may have many millions of files, even a 64-bit hash is not good enough, a hash of 128 bits or more is needed to make negligible the probability of collisions. I have verified this experimentally, finding several 64-bit file hash collisions in my file systems.
- akvadrako 3y agoExcept MD5 is slower than SHA2 on modern PCs / servers.
- Dwedit 3y agoMD5 is incredibly broken. The PDF file PoC||GTFO 0x14 (https://dl.packetstormsecurity.net/mag/pocgtfo/pocorgtfo14.pdf https://dl.packetstormsecurity.net/mag/pocgtfo/pocorgtfo14.p..., 42MB large) is a PDF file that can be also run in a NES emulator, and will display its own MD5 hash. The MD5 hash is also shown in the pdf document itself. (Don't download it from archive.org, their copy is altered) The fact that any document can contain its own MD5 hash embedded in there should be hugely concerning enough. The hash also happens to start with 5EAF00D.
- refulgentis 3y ago[flagged]
- userbinator 3y agoThere's a GIF MD5-quine here: https://news.ycombinator.com/item?id=13823704 https://news.ycombinator.com/item?id=13823704 And a PNG version too: https://news.ycombinator.com/item?id=32956964 https://news.ycombinator.com/item?id=32956964 But no one has made an exclusively plaintext (ASCII) MD5-quine yet, and I suspect doing so may be impossible given the characteristics of collision blocks.
- ForkMeOnTinder 3y agoHow is it impossible? I would think an MD5 quine exists with probability approaching 1 as the size of the document grows to infinity. Think about the reduced problem: 1. a document containing "1", whose hash begins with "1" 2. a document containing "12", whose hash begins with "12" 3. a document containing "123", whose hash begins with "123" #1 is certain to exist. #2 exists, but would take 16x as long to brute force. #3 would take 16x longer again. If this pattern doesn't continue until 2^128, where would it stop, and why? All hashes can be brute forced this way, even secure ones SHA-2. Its security relies on the fact that the earth doesn't contain enough computing power to execute a brute force attack within the universe's lifetime.
- userbinator 3y ago
- userbinator 3y ago[flagged]
- tptacek 3y agoThis would be a better comment without that last line. There's a fair bit of dogma in your comment, too!
- tptacek 3y agoThe unsatisfying answer to this is probably that it just doesn't matter. It's not as if evidence chain of custody is assured cryptographically; it's assured by rules and regulations and an adversarial system. If you tried to submit as evidence a forged document vouchsafed with a colliding MD5 hash, you'd be putting your own freedom at risk, because the forgery will be straightforwardly detectable (the real document won't have hash colliding artifacts in it). None of this is to say that the legal profession shouldn't move to SHA2; it should.
- woodruffw 3y agoThis is more or less what I learned when I worked on forensics software, the kind that was supposed to maintain this kind of chain of custody/integrity. Like most things that touch the legal system, the presumption is that dishonesty or unsoundness in the chain of custody is fundamentally a legal problem with legal recourses, not something that can be solved with math.
- graypegg 3y agoI quite like the licensing trick Nintendo used on the gameboy as an example of this. [0] Essentially, the gameboy expected a bitmap of the Nintendo logo to be present on the cartridge rom, and was shown on screen at boot. It had to match a version stored on the gameboy itself or else the game wouldn’t start. The thinking (that I’m not sure was ever tested) was that someone producing a game that tried to trick consumers into thinking it was an official Nintendo product, would be liable for damages in a trademark lawsuit. Since the game would never start without an official Nintendo logo, the hope was to make the legal system enforce Nintendo’s licensing scheme. [0] https://catskull.net/gameboy-boot-screen-logo.html https://catskull.net/gameboy-boot-screen-logo.html
- deleted 3y ago[deleted]
- bentley 3y agoThe thinking was tested (in U.S. jurisdiction) in Sega v. Accolade. https://en.wikipedia.org/wiki/Sega_v._Accolade https://en.wikipedia.org/wiki/Sega_v._Accolade The court sensibly ruled that using technical means to force competitors to display your trademark against their will doesn’t mean you can then claim they’re infringing that trademark.
- krallja 3y agoI actually picked SHA256 for the path-prefix feature in https://jacob.jkrall.net/benfords-law https://jacob.jkrall.net/benfords-law for “NIST compliance.” That is, I didn’t ever want to answer “yes” to a potential customer’s CISO security surveys question like “does your application use any non-NIST-approved hashing functions?” It’s frankly broken that evidence-handling doesn’t have to follow the government's advice about hash function selection!
- bostik 3y agoFunnily enough, I had an interesting discussion with a client's lawyer (who, to their credit, is reasonably tech-savvy) before the holidays. I had redlined "FIPS 140-2" from their contract language. I'll omit the context, because it's too nuanced to be discussed here, but the long and short of it was that she wanted to know why I did that. I informed her that since FIPS 140-2 is about physical properties of key creation and management, all the relevant layers in a cloud-only solution are simply in the wrong scope. And I added that I am allergic to the string "FIPS" in general. Even having it present in official contract language makes people leap into weird assumptions about supported and allowed algorithms. Her response? "Oh, that makes sense."
- krallja 3y agoMay we all be so lucky to have such an enlightened client(‘ s lawyer)!
- ipython 3y agoI read through this hoping to have a reasonable discussion of the difference between preimage attacks (see https://en.m.wikipedia.org/wiki/Preimage_attack https://en.m.wikipedia.org/wiki/Preimage_attack) and was disappointed when I did not see the topic mentioned once. :( It is much more computationally feasible to create two inputs from scratch that hash to the same value than to forge an existing documents hash (the threat model I’m assuming they’re discussing in relation to the law). As far as I know I am not aware of a demonstrated second preimage attack on md5. Not saying to keep using it, just trying to not spread fud. Edit: I do see second preimage is mentioned about 3/4 of the way through the article. I confess that I did stop reading and started skimming before then.
- userbinator 3y agoYes, if you haven't already noticed, crypto is just a religion at this point, propagated by the "experts" who don't actually think and actively silence dissent.
- evouga 3y agoI think it's simply that the blog author and commentators have an unrealistic threat model when it comes to how the legal profession uses MD5s. After the first high-profile case where authenticity of evidence gets called into question because a seized electronic document was deliberately doctored to allow for a hash collision (if that ever happens), there will be a will to change to something new.
- refulgentis 3y agoI doubt it, legal types won't see this as a math problem[1], but a legal problem (forging documents) [1] unless I'm missing something, this boils down to: "given f(x: string) => y, how can I minimize the odds that you can generate an X for a desired Y"
- deleted 3y ago[deleted]
- adrian_b 3y ago
- kmeisthax 3y agoIf you could just throw anyone who forged a digital signature in prison, you'd keep using MD5, too. The reason why people like us keep changing everything for security is specifically because we have no access to justice. Computer crimes are international and difficult to prosecute, so you might as well drop an algorithm like a hot potato if anyone - even just nation state actors - could break it. We build our rules out of code because we do not have access to the material they make laws out of. That being said, continuing to use MD5 is utterly inexcusable.
- Moru 3y agoYes, this is a thing. My arguments have bene shot down with a handwaving several times. "But that would be a crime so then we call our lawyers". Feels like it would be cheaper to just use something secure than to pay a lawyer :-)
- NoboruWataya 3y agoIf someone commits a crime, it's not the victim of the crime that has to pay for the lawyer to prosecute them.
- m3047 3y agoWould have upvoted except for that last sentence. There is no such thing as a perfectly good airplane, at least as long as "perfect" means flawless rather than good enough. Granted, there is a whole dance for accepting the risk from known defects (that's the whole point).
- throwaway89201 3y agoAnother unfortunately place where MD5 is widely used: pirate libraries such as Library Genesis and Anna's Archive. While content is distributed at large in torrents with SHA1-summed shards, and Anna's Archive at least offers some structured metadata which would allow to slowly migrate away from MD5, files are still indexed using MD5 as primary key, and any other kind of file hash is nowhere to be found. Pirate libraries are particularly important to preserve our cultural heritage in a transparent and trustworthy way. A role that traditional libraries sadly cannot fulfill due to draconian copyright laws, especially around digital books. With archive.org as notable exception.
- bawolff 3y agoIt should be noted that md5 is probably still secure for this usecase (maybe you could do a bait and switch with a specificly prepared file, but you can't force a collision with a non-evil file) Still, they should switch. Sha1 is not good either.
- tesdinger 3y agoI do not need to do a hash collision to upload malware to Library Genesis. I could just upload malware with a slightly different name than a popular book and claim it is a different release, like book_high_quality. To securely view content downloaded from such sites, update your software and sandbox the application.
- callalex 3y agoThis is the same legal system that still uses polygraphs as “lie defectors” and known-junk DNA matching tests as fact, so this isn’t exactly shocking.
- hollerith 3y agoWhat court has admitted polygraph test results into evidence? Surely, none in the US.
- randombits0 3y agoThey are often used in pre-employment screening for sensitive government jobs. Yes, it would be illegal for a private employer to do this.
- callalex 3y agoI suspect we have different definitions of “legal system” in mind. You are correct that such things cannot be admitted as evidence into a court case, but law enforcement agencies still use the machines and do their best to lie to unknowing victims that it will be admitted to court.
- chias 3y agoIt's worth noting that there are no known attacks against MD5 HMACs, which look identical to MD5 hashes.
- notahomosapien 3y agoQuantum computers will severely break MD5 and SHA-1, so they'd be broken even if they are used with HMAC. Use SHA2-256 unless you need quantum-resistant collision resistance, in which case you should use SHA2-384. Use HMAC-SHA2-* with an 256-bit key if you want to prevent length extension attacks.
- bawolff 3y agoSeverely break is a bit of a overstatement. It will make a speed up, but its not like shor's algorithm - you need a really powerful quantum computer before md5 comes under threat. But to be clear. Md5 is broken do not use.
- upofadown 3y agoThe history of this makes it hard to convince people to supersede hashes based on the fact that they can be collided. If the legal community had switched to SHA-1 at the point that MD5 was found to be weak for collisions they would have had to consider switching over to SHA-2 10 years later. From their perspective they dodged a bullet. There ends up being a usability issue here. An MD5 hash is only 128 bits long. So 32 hex digits. A SHA-2 hash is going to be 256 bits. Or 64 hex digits. Manually comparing 64 hex digits is in practice much harder than twice as hard as comparing 32 hex digits. People get lost in the middle. If you chop down your 256 bit hash to 128 bits then due to birthday collisions you can probably brute force a collision anyway (you end up only having to do something like 2^64 operations). So there ends up being a usability argument for specifying that your system has to be able to be secure in the face of collisions. At that point you could then further argue that you will just stick with MD5.
- wuiheerfoj 3y agoRunning 2^64 SHA1 ops on a GPU takes 15 years, so I think finding a reasonable collision for that half using SHA2/3 is not as trivial as you suggest: https://crypto.stackexchange.com/questions/84520/how-long-would-it-take-to-brute-force-a-32-or-16-bit-integer-and-which-type-of-p https://crypto.stackexchange.com/questions/84520/how-long-wo...
- moyix 3y agoSince that post was published, the 4090 came out, which can (according to this hashcat benchmark [1]) do 50,638.7 million SHA1 hashes per second, so now it would only take a single 4090 GPU 11.55 years. Or you could buy 12 of them and do it in a year, etc. So it's definitely not cheap but 15 years is definitely an overestimate (and presumably GPUs will keep getting faster...). SHA2-256 is "only" 21975.5 MH/s so you'd have to double the number of GPUs or amount of time. [1] https://gist.github.com/Chick3nman/32e662a5bb63bc4f51b847bb422222fd https://gist.github.com/Chick3nman/32e662a5bb63bc4f51b847bb4...
- gruez 3y agoGenerating 2^64 hashes isn't guaranteed to produce a collision, and even if a collision did exist in that set, you're not going to find it by getting a bunch of GPUs to compute 2^64 hashes. There's a huge difference between a haystack that maybe contains a needle, and a needle that's been pulled from the haystack and presented to you. To actually find and identify the collisions you'll need to hook those GPUs up to some sort of storage/retrieval system. Just to store 2^64 128-bit hashes would take 295.1 exabytes. That's an order of magnitude more storage than NSA's utah datacenter[1]. [1] https://en.wikipedia.org/wiki/Utah_Data_Center https://en.wikipedia.org/wiki/Utah_Data_Center
- deleted 3y ago[deleted]
- tneely 3y agoWe still see heavy use of MD5 in genomics as well. It's effectively used to generate a single identifier that can be used to reference a specific genome assembly. There have been discussions and attempts to move to other, more secure algorithms, but the community and its tooling is too deeply entrenched in using the MD5 for the reference that it would take a herculean effort to change. I'm personally of the opinion that it doesn't matter. MD5 is fine for genomics. The chances of valid genome files colliding is still extremely low, and there's not really any relevant attack space. Replacing one assembly file with another will just break someone's analysis pipeline, and most likely in a very clear obvious way.
- wyldfire 3y ago> there's not really any relevant attack space. Then why use a cryptographic hash at all? much better hashes out there that only strive for distribution/avalanche. https://en.wikipedia.org/wiki/Non-cryptographic_hash_function https://en.wikipedia.org/wiki/Non-cryptographic_hash_functio...
- drisden84 3y agoMD5 has/had a well-known "media" surface - lawyers/genomics folks had heard of it. Libraries had it as an accessible function (command line utilities, even). Sure, there are better non-cryptographic hashes, but, again the concern of lawyers and genomics folk is neither security nor efficiency - simplicity and "works most of the time" are the two metrics at stake. If either laywers or genomics folks cared about document forgery of this nature (spoiler, they don't), they would move to something like SHA3. If they had a need for high-scalability hash algorithms (spoiler, they don't), they would switch to another faster algorithm. This is a concept I understand security folks struggle to understand - sometimes we _just don't care_. And we never should. Maybe, something a struggling security enthusiast could understand - a video game. If you implement e.g. a caesar cipher, you can have fun, accessible puzzle. Implementing AES in your game as a puzzle, while much harder, fails desperately at the "accessibility" metric. In your single player game, if you want to see some "identifying hash", if you see an md5 one, that's enough. No, you should not worry about people forging documents for your ad-hoc identification system, if you don't have people attempting to forge in-game items. Maybe its even a feature that you need to forge such a hash, as a way to solve a puzzle.
- OhMeadhbh 3y agoThere's a difference between finding a collision and finding a second pre-image. While I agree you shouldn't use MD5, and absolutely don't use a signature algorithm which uses it, finding a second pre-image is harder than finding an arbitrary collision with MD5. An "arbitrary collision" here means you can find two inputs (pre-images) which hash to the same thing. Like you ran some code and discovered that "SDFKLHKLJxchjasdfgklhjaskdhjlf9" hashed to the same thing as "klhkasdfhjkl899078790". Finding a second pre-image means you start with one message, like "ALL QUIET. REMAIN CALM." and figured out that "ATTACK AT DAWN 051928" hashes to the same MIC. I can't believe I'm defending using MD5. But... finding second pre-images is still hard. Sasaki & Aoki say it's got a complexity of around 2^116.9 and requires 11 * 2^45 words of memory (thought 1400Tb isn't THAT outlandish these days.) Still... statements like "finding a second pre-image is hard" don't age well and will guarantee a tractable second pre-image attack will be published tomorrow. But... if you have a bunch of docs and you're not signing them or asking people to trust the hash of each doc, you can (reasonably) quickly de-dup by sorting by MD5 hash and then looking for dups. Which is how many people use MD5. And they continue using MD5 because multiple organizations have similar lists and if you wanted to change it, you would need to get everyone to move to a different algorithm. But yeah... at this point we should assume someone will publish a tractable second pre-image attack "any day now" and get to work migrating from MD5 to MD5 : Next Generation. But good luck getting more than 2 people to agree to what the next preferred hash algorithm should be.
- bawolff 3y agoSo if the document is evidence , then its probably created by the attacker. This seems like a setup where collision is more relavent than 2nd preimage.
- OhMeadhbh 3y agoHow so?
- adrian_b 3y agoSecond preimage attacks are relevant for the documents that you create and give to others. Keeping a hash of the document ensures that you can prove that any altered document shown by someone else is not the original. Collision attacks are relevant for the documents created by others, which you receive. If you have a hash of the document that is collision-resistant, you can trust that the creator does not have other variants of the document with the same hash. If the hash is not collision resistant, i.e. it is MD5 or SHA-1, you cannot know if the creator of the document has not also created another variant of the document than the one handed to you, which has the same hash. That is why a digital signature on a document received from others is meaningful only if it is based on a collision-resistant hash. If you sign and verify your own documents, for detecting modifications, a second preimage attack resistant hash would be enough.
- tgamblin 3y agoThe article mentions the key detail: MD5 is broken for cryptography (collisions) but not for second preimage attacks. I was hoping there would be some discussion of just how much more difficult the latter is. It is extremely difficult. Let’s ignore that no second preimage attack is currently known for MD5. The software the author links to has a FAQ that links to a paper that lays out the second preimage complexity for MD4: https://who.paris.inria.fr/Gaetan.Leurent/files/MD4_FSE08.pdf https://who.paris.inria.fr/Gaetan.Leurent/files/MD4_FSE08.pd... It takes 2^102 hashes to brute force this for MD4, which is weaker than MD5. A bitcoin Antminer K7 will set you back $2,000, and it gets 58 TH/s for sha256, which is slower than MD5 or MD4. Let’s ignore that MD5 is more complex than MD4, and let’s say conservatively that similar hardware might be twice as fast for MD5 (SHA256 is really only 20-30% slower on a cpu). It’ll take 2^102/58e12/2/60/60/24/365, or about 1.4 billion years to do a second preimage attack with current hardware. So you could do that 3 times before the sun dies. If you want to reduce that to 1.4 years, you could maybe buy a billion K7’s for $2 trillion. And each requires 2.8kW so you’ll need to find 2.8 terawatts somewhere. That’s 34 trillion kWh for 1.4 years. US yearly energy consumption is 4 trillion kWh. It will be a while, probably decades or more, before there’s a tractable second preimage attack here. Yes, there are stronger hashes out there than MD5, but for file verification (which is what it’s being used for) it’s fine. Safe, even. The legal folks should probably switch someday, and it’ll probably be convenient to do so since many crypto libraries won’t even let you use MD5 unless you pass a “not for security” argument. But there’s no crisis. They can take their time.
- gojomo 3y agoSecond preimage attacks aren't the only threat in a forensics environment. Also, hand-wavy extrapolations from Bitcoin miners aren't a reliable estimate of how fast & energy-efficient dedicated MD5 hardware could become.
- tgamblin 3y agoWhich part was hand-wavy/unreasonable? Do you think that dedicated MD5 hardware could become billions or even millions of times more efficient within a decade? If so, why?
- bawolff 3y agoI've always wondered what would happen if some black hat made all their payloads have the same hash in the hopes that if it ever went to court the confusion would hinder the proceedings. Like i get that md5 is essentially a unique identifier and not meant to protect against malicious interference but if all the exhibits had the same identifier surely that would confuse people.
- emmelaich 3y agoInterestingly, use of MD5 has been been found in court to make evidence less convincing. https://www.schneier.com/blog/archives/2005/08/the_md5_defense.html https://www.schneier.com/blog/archives/2005/08/the_md5_defen... To the article, what tptacek said.
- StillBored 3y agoSigh, been having this conversation in a related codebase. Md5 is just as fine as any other generic hash function if its being used as a non-unique key, which for many cases replacing it with one of the more "secure" alternatives does nothing except for the fact that the resulting hashes are frequently longer, thereby further reducing the statistical chance of an accidental collision. For something like a document store, duplication system, etc, simply taking the extra step of doing a binary comparison against the text associated with the hash assures that accidental (or intentional) collisions are handled. With the bonus that you probably get to either publish a paper or detect someone trying to attack the system should the text comparison fail. And given the history of cryptographic hashes, i'm even more convinced that anyone depending on sha3/whatever being better than md5/etc over the next 10-20 years is fooling themselves. Now would I use it in a secure boot chain/etc as a stamp of uniqueness? Probably not.
- jochem9 3y agoThe chances of having an accidental hash collision are really small. I have build data warehouses with md5 as the hashing algorithm to generate keys from natural keys. Did some back of the envelope calculations back then and found that the chance of a hash collision was minute. Don't remember the exact numbers, but somewhere in the 100s of years if I was generating keys every second. This could btw very well be a thing with large volumes of data, but in many systems this absolutely not a worry.
- brohee 3y agoThe chance of a random collision is minute but if someone is actually building collisions the system is broken. DVC uses MD5 of file for reduplication for example and when you purposely inject files withe the same MD5 (which take seconds to build) the result is data loss.
- Moru 3y agomd5 is faster due to being older and made for older hardware so I guess that is why it's in use for things like that. All deduplicating tools I have used first check for file length before it even tries to do a checksum so I guess that would take care of some problems. It's harder to find a collision if you have to keep the filesize the same.
- CaliforniaKarl 3y agoI've been wondering, is there a term for a type of attack like this: Given a message M, length function L(), and MD5 hash function H(); is there an attack which can generate message M', such that H(M)==H(M') _and_ L(M)==L(M')? In other words: Two different messages, both of the same length, with the same hash? It's almost like a chosen prefix collision attack, but with no prefix (so P is empty) and a given message (M is known, M' is up to the attacker). I ask because I frequently use GridFTP for data transfer, and it uses both the file length and the MD5 has to verify that files were transferred correctly.
- supriyo-biswas 3y agoThat is still an attack on the second preimage or a collision resistance properties of the hash function. Most collisions do work this way, for example see [1]. [1] https://github.com/corkami/collisions https://github.com/corkami/collisions
- CaliforniaKarl 3y agoThat makes sense, but is there a specific name for this type of collision?
- abhibeckert 3y agoI don't know anything about GridFTP - but there's a huge difference between verifying if files were "transferred correctly" and verifying that files were transferred without being tampered with by a malicious party. MD5 is fine for the first task, and totally unacceptable for the second.
- CaliforniaKarl 3y agoIndeed, which is why I didn’t mention third-party tampering. For that, the transfer can be sent inside of a TLS-enabled connection.
- marcus_holmes 3y agoI worked for years as the tech guy for a document imaging company. We worked a few gigs for the Serious Fraud Office in the UK. So, while I'm not a lawyer, I bumped into this stuff a fair bit. The point is that evidence is an agreement between the two sides in a case, and it's not an absolute thing. If you have the original document that was signed by both parties, great. If you have a scan (using a lossless compression format) of the document and proof that the original was destroyed, great. If you have a scan but no proof of destruction, still great. If you have a photograph of the document and no proof, still great. If you have a vague recollection of what was in the document, still great. All of these are "great" if the other side accepts that they are accurate depictions of the original. If they don't accept that, then there's an argument about what the original document contained and the provenance of the evidence, and only then does the actual quality matter. Original document with wet signature is hard to argue with (but not impossible - wet signatures can be forged). The further away from that, the easier it is to argue that the document presented is not accurate and should not be accepted as evidence. Knowing that it's possible to use collisions to create false evidence doesn't matter if no-one contests that the evidence is false. It only becomes significant if one side says that the document has been tampered with, and that's not that common. The side claiming it was tampered with would have to present their version of the document, and their version of events that allowed the document to be tampered with, and so on. The judge would make a ruling about which version of the document was considered the "real" one and the case would continue. Obviously there are edge cases where the whole trial verdict hinges on which version of the document is the correct one, but they're edge cases. And in those cases you could-re-hash the documents involved and double-check with one was right, etc. In the OP's example, where a letter of recommendation has the same hash as a authorisation letter, this is only going to matter if one side says the accused was authorised and the other says they weren't. The authorisation letter will be produced by one side, and the recommendation letter produced by the other, and there'll be an argument about which was the original document. The fact that they have the same hash isn't really relevant. It's a minor point of interest given that these are two clearly different documents saying different things. In the specific cases for the SFO that I worked on, the SFO descended on the accused's offices like locusts, sweeping every single document into carefully numbered bags. We scanned the documents in secure facilities, stored the originals in secure facilities, stored the resulting images in secure storage, and deleted any cache or copies. My professional opinion is that it would be impossible for anyone to create two documents prior to the SFO's investigation that would create an intentional MD5 collision in the evidence used in court. And, even if they somehow did, it wouldn't matter because both documents would be in evidence bags in storage and could be recovered to be examined by the court. Obviously, from a black/white technical point of view, using a better hash algorithm would be better. But I can see why the legal profession is reluctant to adopt the new thing; it's a hassle and it will only affect a tiny amount of cases, if any.
- wglb 3y agoDespite official statements about MD5, the leading e-discovery software provider uses sha256: https://help.relativity.com/9.0/Content/Relativity/Processing/De-duplication_considerations.htm#:~:text=Relativity%20calculates%20the%20SHA256%20hash,loose%20files%20to%20identify%20duplicates https://help.relativity.com/9.0/Content/Relativity/Processin.... And it is unclear if that is in any way unusual.
- tcper 3y agoIn other industries, MD5 collision is the most minor problem, it just works, and God know when/how they encounter collision, so just use it.
- xpil 3y agoInterestingly, DBT is still using MD5 as the default algorithm for generating row identifiers.
- robertlagrant 3y ago> Yes, they say, MD5 is broken for encryption, but since they’re not doing encryption, it’s fine for them to use it. Unless I missed it, this article seems to not refute the most fundamental point: MD5 was never broken for encryption. Hashing is not encryption.
- adrian_b 3y agoWhile hashing is not encryption, any secure hashing function can be used for encryption, even when used as a black box, (by making an unpredictable PRNG with it). Moreover, MD5, SHA-1 and SHA-2 contain a block cipher function used in the Davies-Meyer mode of operation. The internal block cipher function can be extracted and used in any other mode of operation possible for block cipher functions. Because of these possibilities, many older laws that have existed in various places, prohibiting the inclusion of encryption in software products, but allowing secure hashing functions, have been completely misguided.
- alextingle 3y agoIt's broken for hashing too. The point is that MD5 is no good if there's any way an adversary might want to subvert it. It's fine if you just want to use it for hashing your own documents, but as soon as there's an incentive for someone to substitute one document for another, MD5 is problematic. That's certainly the case for encryption, but it's also the case for these legal document records.
- robertlagrant 3y agoI don't disagree, but I think my point still stands.
- heads 3y agoI can see the lawyers’ point. Any old checksum is good enough for spotting random data corruption. If you do happen to have a crooked lawyer who is submitting tampered evidence then you would hope that there are better systems in place to weed out this ethical corruption. For instance: law school training to strengthen ethics, well known punishments to act as a deterrent, hiring processes to filter out unscrupulous actors, and whistleblower protections to encourage and reward vigilance. It’s not an opinion that adheres to the cynical zeitgeist, but in my experience most members of this profession are extremely trustworthy. I’m sure they dislike the stereotype of lawyer=rotter just as much as hackers are tired of being typecast as Newman… sorry, Dennis from Jurassic Park!
- laserbeam 3y ago> MD5 should be considered broken and unsuitable for further use. Ya know... It's 2024 and Azure's blob storage ONLY supports MD5 for integrity checks when writing blobs. There are no other hash functions supported there. The default cloud storage solution implemented by one of the largest cloud providers out there ONLY uses MD5. I really want to use something else, but whenever I have to interact with them I must fall back to MD5. It's not up to me as a dev to use something better if I need to interact with Azure. Yes, I can use other hashes alongside MD5, but if I want integrity checks with the storage provider I can't completely abandon MD5.
- Pikamander2 3y agoWordPress still uses MD5 for database passwords to this very day with no immediate plans to change it. That said, they apparently use eight passes of MD5 hashing along with salting, which they claim is a sufficiently secure combo. WordPress's core and default themes are known to be fairly secure, so I'd like to believe they know what they're talking about, but if nothing else it feels icky.
- Snow_Falls 3y agoI'm confused, if they're going through the effort to make something known bad (MD5 secure, then why not just use something secure in the first place (e.g. SHA3)?
- shpx 3y agoMD5 hashes are half the length of the recommended hashing algorithms. This convenience (and the switching cost) is worth more than the theoretical security considerations.
- NoboruWataya 3y ago> (and the switching cost) Yes, I must say I smiled when I saw the author's assertion that moving the entire legal industry to a new hashing algorithm is "trivial".
- denton-scratch 3y ago> only broken for encryption It's broken in an adversarial situation: given the hash of evidence-file A, it's possible to construct a file B that gives the same hash. But it would be a different matter entirely to construct a file B that actually looked like a file of evidence relevant to the case. I don't know how lawyers use these hashes, but unless they're being used to detect malicious tampering, I don't see what's wrong with MD5. And since the files to be hashed are evidence, they're in the custody of a court; things have got quite bad if court officials might be tampering with evidence.
- aljarry 3y ago> It's broken in an adversarial situation: given the hash of evidence-file A, it's possible to construct a file B that gives the same hash. No, that's a second preimage attack. MD5 is safe against preimage & second preimage attacks. What MD5 is not safe against, is a collision attack: you can create two messages/files with different content, that end up having the same hash.
- denton-scratch 3y agoYeah, sorry. TFA made that clear. So to exploit the vulnerability, you have to be able to manipulate file A, the original piece of evidence, to construct a file B that has a matching hash. I still fail to see how this impacts files submitted to a court in evidence.
- LgLasagnaModel 3y ago“A cryptographic hash should uniquely identify a file.” This is not true. It’s not possible to guarantee this. One can be certain that two files DON’T match if they have different hashes, but one cannot be certain that two files DO match based ONLY on the fact that they have the same hash.
- ddtaylor 3y agoThis is actually something I discussed with my legal team and was prepared to bring up at trial with our own forensic expert. My team was (correctly, IMO) not expecting much to happen because the amount of precedent that exists with MD5 being used as a way to say these documents haven't been swapped or tampered.