3 ms·
If you look at the assigned identifiers in the binary scheme, the authors believe that 32-bit truncated cryptographic hash functions are useful. No they don't.
by davidp 11y ago
If you look at the assigned identifiers in the binary scheme, the authors believe that 32-bit truncated cryptographic hash functions are useful.
No they don't. From Section 2:
The sha-256 algorithm as specified in [SHA-256] is mandatory to
implement; that is, implementations MUST be able to generate/send and
to accept/process names based on a sha-256 hash. However,
implementations MAY support additional hash algorithms and MAY use
those for specific names, for example, in a constrained environment
where sha-256 is non-optimal or where truncated names are needed to
fit into corresponding protocols (when a higher collision probability
can be tolerated).
Truncated hashes MAY be supported. When a hash value is truncated,
the name MUST indicate this. Therefore, we use different hash
algorithm strings in these cases, such as sha-256-32 for a 32-bit
truncation of a sha-256 output. A 32-bit truncated hash is
essentially useless for security in almost all cases but might be
useful for naming. With current best practices [RFC3766], very few,
if any, applications making use of names with less than 100-bit
hashes will have useful security properties.
- btrask 11y ago32-bit hashes are not useful for content addressing pretty much at all. If you look at the table here: https://en.wikipedia.org/wiki/Birthday_problem#Probability_table https://en.wikipedia.org/wiki/Birthday_problem#Probability_t... (50% chance of collision with only 77,000 resources) See also: https://lkml.org/lkml/2010/10/28/287 https://lkml.org/lkml/2010/10/28/287 where Linus Torvalds says 12 hex digits (96 bits) is pretty much the minimum short-hash for the Linux kernel commit history. Extremely short hashes can be useful briefly for manually transcribing between devices, as long as you immediately "resolve" them back into a longer form, before new collisions can happen. But this is more on par with clicking "I'm feeling lucky" than creating a link. :)
- batbomb 11y agoI came here to talk about this. I deal with data management for a bunch of physics and astronomy experiments it's quite typical to have many millions of files for any given experiment (say, 1PB of storage for a medium size experiment, 10 million 100MB files). CERN's experiments would easily have billions of files, so I was thinking 64 bits would be a minimal truncation, but 96 is probably more reasonable.