4 ms·
Can you explain what it means to "check the data"? The point here is that, for the last year, Twitter has NOT been generating sequential IDs. Just look at the s
by joshma 14y ago
Can you explain what it means to "check the data"? The point here is that, for the last year, Twitter has NOT been generating sequential IDs. Just look at the source code[1]. They save a few low bits for the worker ID, data center, and worker-specific sequence. The highest bits are the timestamp bits.
Maybe you can correct me if my assumptions based on the code are wrong, but my guess is you'll magically find that Twitter's highest ID grows linearly with time. Then you're just randomly sampling within this maximum timestamp. At this point you're assuming that Twitter sees a constant rate of signups wrt time, which I highly doubt.
[1] https://github.com/twitter/snowflake/blob/master/src/main/scala/com/twitter/service/snowflake/IdWorker.scala https://github.com/twitter/snowflake/blob/master/src/main/sc...
- diego 14y agoWhat I did was to measure the difference in yield between uniformly generated ids that would correspond to the time after Snowflake (when ids were at 380M) and the ones before. It's true that the yield is less. It went down from about 86% to 82%. A separate problem is that my estimate of the highest id at the time of the experiment likely fell short. Since then I've encountered higher ids that are pretty sparse, but I don't know how many there are. Do you have any ideas as to how to generate a better sequence of random ids that tracks Twitter ids after Snowflake? I'd like to redo this experiment in a while.