7 ms·
Guid Smash
- nesk_ 1y agoNice experiment. Is the code available somewhere?
- 8organicbits 1y agoNote that this only considers UUIDv4, the random UUID. Other forms can generate UUIDs that are much closer together. For UUIDv7, UUIDs generated within the same millisecond will have identical 48 bit prefixes (or up to 60 when the monotonic counter from section 6.2 is used). https://www.rfc-editor.org/rfc/rfc9562.html#monotonicity_counters https://www.rfc-editor.org/rfc/rfc9562.html#monotonicity_cou...
- e1g 1y agoYou need to be generating >100M of them within the same millisecond before even remembering that collisions can theoretically happen.
- charcircuit 1y ago>You The entire universe. Else it's not universally unique.
- 8organicbits 1y agoI like UUIDv7s as database IDs since they sort chronologically, are unique, and are efficient to generate. My system chooses the UUIDs; I don't allow externally generated IDs in. If I did, then an attacker could easily force a collision. As such, I only care about how fast I create IDs. This is a common pattern. If your system does need to worry about UUIDv7s generated by the rest of the universe, you likely also need to worry about maliciously created IDs, software bugs, clocks that reset to unix epoch, etc. I worry about those more than a bonefide collision.
- tonyhart7 1y agoYour app is must be popular to be having an entire universe "amount" of users lol joke aside all of this is theorical, in practical application its literally impossible to hit it that it doesn't matters if its possible or not since you are not google scale anyway
- charcircuit 1y agoIt's not just your app. It's any other app or data provider that you may now or in the future interact with.
- tgv 1y agoOnly if the other side uses your key as theirs, and uses it to store data from many sources. I, personally, don't feel it's hardly worth considering. A primary key under your own control doesn't cost much, and is a better choice.
- Xelynega 1y agoThat's not how namespacing works though, is it? Getting UUID 'A' from app 'X' is easily distinguishable from UUID 'A' from app 'Y'.
- charcircuit 1y agoThe point of the first U in UUID, universal, is that you don't need to use namespacing.
- tonyhart7 1y agoUniversal mean unique that uid wouldn't be used anyone else in any point in history or just universal available in one app???? because you just overreach at this point, if you can develop a better one. be my guest
- senko 1y agoObviously, just the part within our light cone.
- sgentle 1y agoApparently there's 500 hours of video uploaded to YouTube every minute (30 seconds every millisecond). Assuming 4K@60fps, that works out to 14,929,920,000 pixels per millisecond. If YouTube wanted to give every incoming pixel its own UUIDv7, they'd see a collision rate just under 0.6%.
- avar 1y ago> Assuming 4K@60fps [...] they'd see a collision rate just under 0.6% This doesn't detract from your point of collisions like that being viable at that scale, but assuming an average of 4K@60fps is assuming a lot. The average video upload there is probably south of 1080p@30fps.
- Xelynega 1y agoYou're glossing over the fact that they assumed youtube would want to assign a UUID to each pixel in a 4k@60fps video as the use case that this would fail for...
- e1g 1y agoExcellent example. And at that scale, you are generating 100TB/s in UUIDs so if you need to store them, you have much bigger problems than collisions.
- amingilani 1y agoInstead of picking a target UUID and evaluating new UUIDs against it, a better experiment would be finding duplicates in all the UUIDs you have generated. This plays nicely with the birthday paradox.
- whyever 1y agoIt would require a lot more memory, because you have to remember every generated UUID. And how would you do the partial match? You are not going to observe any collisions.
- twiss 1y ago> The chances of generating two GUIDs that are the same is astronomically small. > The odds are 1 in 2^122 — that’s approximately 1 in 5,000,000,000,000,000,000,000,000,000,000,000,00. This is true if you only generate two GUIDs, but if you generate very many GUIDs, the chance of generating two identical ones between any of them increases. E.g. if you generate 2^61 GUIDs, you have about a 1 in 2 chance of a collision, due to the birthday paradox. 2^61 is still a very large number of course, but much more feasible to reach than 2^122 when doing a collision attack. This is the reason that cryptographic hashes are typically 256 bits or more (to make the cost of collision attacks >= 2^128).
- Retr0id 1y ago2^61 isn't even that large, well within the compute budget of mere mortals.
- PaulHoule 1y agoI think you might have trouble if you tried to assign one to every iron atom in an iron filing.
- vlovich123 1y agoDepends on what “isn’t even that large means”. A modern 6ghz machine would probably need 12 years of 24/7 operation to count that high. To me that seems like a lot.
- dgrin91 1y agoYeah, but a nation state server farm can probably cut that down to minutes because their budget can buy a lot of processors. You only need a few hundred to really shrink it down to manageable numbers. And it turns out that nation starts aren't the only ones that have this budget
- 8organicbits 1y agoWhat's the threat here? It's trivial to force a collision. Here's the same UUID twice: 6e197264-d14b-44df-af98-39aac5681791 6e197264-d14b-44df-af98-39aac5681791 Typically, you don't care about UUIDs that aren't in your system and you generate those yourself to avoid maliciously generated collisions. Your system can't handle 2^61 IDs. It doesn't have the processing power, storage, or bandwidth for that to happen. Not to mention traditional rate limiting.
- webstrand 1y agoThis is the chance that given a specific guid, that you'll find a collision for it. Utterly minuscule chance. However birthday paradox controls, if you generate 2^62.60 guids the chance that you've generated a collision is around 99%. Still enormously unlikely, but way smaller than 2^122. At a rate of comparing 400,000 guids per second, you have a 99% chance of seeing a collision within the next 553,750 years.
- jonathrg 1y agoYou would need a little more memory to see/detect that collision.
- RS-232 1y agoUUID > GUID. Microsoft’s GUID standard is garbage.
- lionkor 1y agoOh, why?
- w-ll 1y agonot OP but i already have fields for time ts and what model it is. i want my uuids random.
- kaoD 1y agoI think the current Microsoft GUID is just UUIDv7. https://learn.microsoft.com/en-us/dotnet/api/system.guid?view=net-9.0 https://learn.microsoft.com/en-us/dotnet/api/system.guid?vie... I don't think there's a "Microsoft standard" and they just use different versions of UUID in different products over time. No idea why they call it GUID instead of UUID though, but it's easier to speak out loud so I'm not against it. v7 has a timestamp indeed, but isn't the time making it more collision resistant? You'd have to generate tons of UUIDv7s in the same millisecond, while v4 is more likely to collide due to not being time-constrained and the birthday paradox. I think both have their uses though. You might need pure random if you want your UUID not to convey any time information and you're not generating tons of them (e.g. a random user id). What do you mean "model"? Are you referring to UUIDv1 which has time and MAC address?
- Zambyte 1y ago> isn't the time making it more collision resistant? That seems to depend a whole lot on the pattern your application generates UUIDs in. If you're generating a consistent distribution over time, sure. If you generate a whole lot in bursts, collision seems to be way more likely.
- kaoD 1y ago
- curtisszmania 1y ago[dead]
- nopassrecover 1y agoReminds me of a problem I ran into once where someone had wanted unique but short codes as identifiers for relatively small counts, and picked a substring of a UUID: http://mattmitchell.com.au/birthday-problems-friendly-identifiers-and-mongodb/ http://mattmitchell.com.au/birthday-problems-friendly-identi...
- kr2 1y ago> However, the overall takeaway was: Don’t use the MongoDB Increment value as a Unique Identifier. However, the overall takeaway should be, as always: don't use MongoDB. Period. Every time I learn something new about it I'm baffled about why people continue to use it.
- Joel_Mckay 1y agoMost just pack down: epoch time + MAC Address + transaction counter (catch NTP skew) + Thread PID + new Pointer address = GUID Then increment global transaction counter, complete some ops, and check to ensure current epoch time is in the future before the transaction frees the memory locations. This is often robust in highly concurrent distributed systems even under network degradation, or corrupted sync states. Has other interesting use-cases too. =3
- ahmedfromtunis 1y agoThe proximity measure seems to be flawed. If you want to see how close to a non-ordinal 123456 a random generator can get, you also need to look for stuff like 923456 or 123956, etc. Also, would 223456 be considered a closer match compared to 323456? (It shouldn't in my opinion because, again, these are non-ordinal strings).
- gammalost 1y agoIf its a random ID then I'd argue that all of them are equally close to each other. With that said, I do not know how GUIDs are generated
- franky47 1y agoEasy, it should be listed there: https://everyuuid.com/ https://everyuuid.com/
- 867-5309 1y agoplease may all the death huggers go hug a tree. thanks
- ivanjermakov 1y agoReminds me of SHAllenge: https://news.ycombinator.com/item?id=40683564 https://news.ycombinator.com/item?id=40683564