8 ms·
Show HN: Sonyflake – distributed unique ID generator implemented in Rust
- abahlo 6y agoNeeded this functionality in a Rust application I'm writing so I ported the Go code to Rust.
- echelon 6y agoNice work! I'm curious, if you don't mind my asking. What are you building that requires that amount of scale? These are heavy hitting IDs (but of course you know that).
- social_quotient 6y agoMaybe they need the distribution part and not the scale. But I’m curious now too.
- abahlo 6y agoThanks and I don't mind at all! To be honest for the application I didn't want to expose a database serial (otherwise you'd know you're user #100, for example) or use UUIDs, so it's less about scale and more about obscuring ids. The library is well suited for huge scale scenarios nonetheless.
- LaundroMat 6y agoWhy did you not want to use UUID's?
- abahlo 6y agoI already use UUIDs for various fields and didn't want the id to be confused with other fields. More of a clarity/style decision than technical.
- tinus_hn 6y agoUnfortunately this means you get to make the same mistake as originally made with UUID: the static machine ID is a privacy leak.
- abahlo 6y agoIs it? Not sure private ips are so sensitive, esp. the lower parts.
- tinus_hn 6y agoIf it’s unique it’s a privacy leak. I’m not going to debate that, it’s common knowledge. UUID was changed in the 90s to be a hash of this instead and later to just be a completely random number because there are so many bits the likelihood of a single duplicate being generated before the sun has swollen enough to consume the earth is slim so you don’t actually need these schemes to provide a unique number.
- deleted 6y ago[deleted]
- bojanz 6y agoWhat are the benefits of Snowflake compared to ULID, which already has Rust implementations[1][2]? [1] https://crates.io/crates/ulid https://crates.io/crates/ulid [2] https://crates.io/crates/rusty_ulid https://crates.io/crates/rusty_ulid
- anomaloustho 6y agoNot sure if this is also a limitation of Snowflake. But ULIDs are only unique down to a certain timescale. (milliseconds I believe)
- sylvain_kerkour 6y agoNot exactly, They can be sorted, by default, only down to a millisecond, but you can use a monotonic generator to have them sorted, even if more than one Ulid is generated within a millisecond. Other than that, they have 80 bits of randomness, enough to be unique even if millions are generated per second. https://github.com/ulid/spec https://github.com/ulid/spec
- abahlo 6y agoSonyflakes are generally more useful in massive scale environments as you won't run into conflicts due to the "machine-id" part (which defaults to the lower 16 bits of the private IP address). There are more differences (128bit vs 64bit, rng vs time-based), I think they have different use cases.
- eutectic 6y agoYou shouldn't ever have a collision with 128 bits of entropy.
- j-pb 6y agoMore structure actually increases the chance for collisions. Because at that scale the chance of bit errors is dominating the equation. More rng protects against that.
- ComputerGuru 6y ago
- deleted 6y ago[deleted]
- random5634 6y agoHow do you use 64 bits and only get scale to 180 years? For that bit count I’d expect 1000+
- abahlo 6y agoA Sonyflake ID is composed of 39 bits for time in units of 10 msec 8 bits for a sequence number 16 bits for a machine id
- olalonde 6y agoWhy do the bits add up to 63 rather than 64?
- sjansen 6y agoJust a guess without reading the code, the unused bit is probably to avoid complications from signed integers. If the sign bit is never set, signed and unsigned integers can store the same data.
- lmilcin 6y agoPresumably to make them work with languages that don't recognize unsigned integers? For example Java...
- secondcoming 6y agoDoes this mean that using them in a multi-threaded context may produce duplicate ids? Edit: To answer my own question. No, although it does seem to limit your application to a single process with at most 256 threads, if I'm reading it correctly. > Sequence numbers are per-thread and worker numbers are chosen at startup via zookeeper (though that’s overridable via a config file). [0] [0] https://blog.twitter.com/engineering/en_us/a/2010/announcing-snowflake.html https://blog.twitter.com/engineering/en_us/a/2010/announcing...
- random5634 6y agoThis proves my point. 39 bits for time in 10 msec is something like 17x years. ULID runs through 10889 AD
- sylvain_kerkour 6y agoAfter a lot of time spent investigating the different kind of (U)UIDs, I've come to the conclusion that ULIDs[0] are the best kind of (U)UIDs you can use for your application today: * they are sortable so pagination is fast and easy. * they can be represented as UUID in database (at least in Postgres). * they are kind of serial, so insertion into indexes is rather good, as opposed to completely random UUIDs. * 48 bits timestamp gives enough space for the next 9000 years. [0] https://github.com/ulid/spec https://github.com/ulid/spec
- minitoar 6y agoNot good if you’re using them in a situation where they are visible to the public (eg urls) and you don’t want to leak the time component.
- andy_ppp 6y agoIf you really need this why not add a public ID...
- mnahkies 6y agoTrue, but for many entities the time they were created is basically public information anyway - time of a tweet, date hacker news account was created, time blog post was published. It might leak more granular information than would be publicly available otherwise, but for many use cases it's not a security concern as far as I can tell
- lmilcin 6y agoUnfortunately, 8 bits for sequence for each 1/100th of a second is way too low. I gather the person that designed this has never seen high traffic service. But for some having cap of 25.6k unique ids per second on a node is a deal breaker.
- Jweb_Guru 6y agoThere are very, very few services which actually need to handle more than 25.6k unique IDs per node per second. I'm not saying they don't exist, but it is not a problem most companies have.
- lmilcin 6y agoAh, yes... and 640kB of RAM was supposed to be enough, too. Or 4B IP addresses. If anything, my experience is whenever somebody sets close hard limits like that somebody else is going to pay for it. The service I work on creates up to couple millions of unique objects per second, on a single node in a batch process. It tries to do a large batch of work as fast as is possible. It fetches I think something like 2GB of data per second from MongoDB node, processes it and creates up to 5 million objects in span of 2-3 seconds. We also have realtime services that are about 20-30k TPS per node. We are currently using MongoDB and it uses plenty bits for that: https://docs.mongodb.com/manual/reference/method/ObjectId/ https://docs.mongodb.com/manual/reference/method/ObjectId/
- secondcoming 6y agoWhy the hell has this been downvoted? Not everyone on HN is a web developer. Some of the audience works on high-throughput systems. Instead of downvoting, perhaps people would prefer to learn? They may save their company lots of money!
- lmilcin 6y agoHey, it is a Rust library. So probably not many web devs are going to use it. Which is even more reason to point out limitations like that because if you choose Rust for your project it means very likely you have some important performance constraints. And that is either linked with a small system with low resources (in which case it is ok) or a large system with huge traffic (in which case it is not).
- Dowwie 6y agoBefore starting this, did you look at the flaken crate? If so, any concerns or notable differences? https://github.com/bfrog/flaken https://github.com/bfrog/flaken Although it hasn't been maintained in 4-5 years, it's not the kind of library that requires any maintenance following creation of a stable generator. Maybe someone could bench it.
- jhgg 6y agoDiscord uses snowflake based IDs to great success. It's actually been one of the rock solid parts of our infrastructure and the fact that time stamps are embedded in the ID are great. Every channel, message, stream, attachment, user, server, etc... gets a snowflake. We originally wrote this service in Elixir but have since rewritten it in rust. Actually very similar to this crate. Maybe I should open source ours too haha.
- Dowwie 6y agoDid you see the flake crate?
- jhgg 6y agoThe one that says "wip; use at your own risk" with no updates in the last 5 years? Or is there another? https://crates.io/crates/flake https://crates.io/crates/flake
- Dowwie 6y agosorry, missed the N at the end -- I refer to "flaken": https://crates.io/crates/flaken https://crates.io/crates/flaken snowflake libraries generally aren't updated after they're stabilized
- jhgg 6y agoNo I didn't but after looking at it it looks like its implementation falls short. - it only looks at system time at system startup, then relies on instant for monotonic clock source. unfortunately for setups that do leap-second smearing (ours does), this would deviate slightly out of sync after operating for a while. - does not handle seq overflow within a given millisecond window - this i think is very unsafe, and can cause the generation of duplicate IDs if `next_id` is called at too rapid an interval.