5 ms·
Instagram uses basic base62 with a very standard alphabet ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789-_ They started using hashing functi
by doh 7y ago
Instagram uses basic base62 with a very standard alphabet
ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789-_
They started using hashing functions from Facebook after they got acquired [0]. This function is used across the whole FB and the main benefit is that it's time sensitive. That means they will sort based on time.
First few millions of IDs are just autoincrements and then they swapped to the custom sort. If you are bored, you can dig through who used instagram in the early days.
BTW base(62/63/64) is very popular way to encode IDs. Some sites will use much more custom alphabet. My favorite was Vine which used
BuzaW7ZmKAqbhMOei5J1nvr6gXHwdpDjITtFUPxQ20E9VY3Ll
[0] https://instagram-engineering.com/sharding-ids-at-instagram-1cf5a71e5a5c https://instagram-engineering.com/sharding-ids-at-instagram-...
- rbranson 7y agoThe ID generation code doesn’t use hashing and has nothing to do with Facebook. The blog post was published afterwards but the sharding code predated the acquisition.
- doh 7y agoMaybe I wasn’t clear. The integers that enter the base62 used to be autoincrements. After time they became what was published in the blog post. If “timestamp” function is not from FB or not inspired by, then it must be a massive coincidence because FB’s IDs are almost exactly the same. They just don’t hash them with base62.
- rbranson 7y agoFBIDs are complicated because the number space includes legacy IDs, but the modern space have a 24-bit shard ID prefix and a 38-bit autoincrement identifier. Two of the bits are “thrown away.” Timestamps aren’t involved. EDIT: ugh. Still early for us slackers on the West coast to do basic math.
- doh 7y agoFair enough. They do appear to be very similar in structure and they do follow some logic in correlation to timestamp. However you seem to be more knowledgeable about it and so I will add this to my mental knowledge book. Thank you
- blotter_paper 7y agoWhat do you like about the Vine ID alphabet? I don't see the appeal, it cuts out some characters yet still includes potentially confusing ones (I/l/1 and O/0, which can look identical or nearly identical in some fonts). If not human readability, why did they cut out standard characters thus making IDs longer than necessary? I suspect there is a perfectly reasonable explanation, but I'm missing it.
- doh 7y agoMy guess is that they used this as a security through obscurity [0]. They hoped that nobody will be able to reverse back the alphabet if they don't get their hands on real IDs. Because we [1] crawl every video/audio site, we had all the links for vine. I just took 1,000 consequential IDs (based on creation time) and then worked backwards on the alphabet. Took maybe 10 minutes to write the code. [0] https://en.wikipedia.org/wiki/Security_through_obscurity https://en.wikipedia.org/wiki/Security_through_obscurity [1] https://pex.com https://pex.com
- trenning 7y agoSince it's not listed on the site, what are the costs for the services pex offers?
- doh 7y agoDepends on the service. We are currently closed to only certain customers, although our attribution engine [0] is completely free to _all_ rightsholders. [0] https://blog.pex.com/introducing-the-attribution-engine-19eccd3441a9 https://blog.pex.com/introducing-the-attribution-engine-19ec...
- crazygringo 7y agoTo me the randomized characters are a clever way of hiding "autoincrement" to users, so it would be non-trivial to calculate the number being produced daily. Picking base 49 (a seemingly random number) seems to be perhaps a similar security-through-obscurity thing? I've written code to produce random (non-sequential) ID's for storage in a database, but it's surprisingly non-trivial to produce a performant INSERT that will work without ever producing collisions. Relying on the database's AUTOINCREMENT but obfuscating the value through a baseXX encoding seems actually much easier, when the only "security" you're trying to provide is against reporters and business analysts trying to estimate the site's usage/popularity, and they're quite unlikely to bother to reverse-engineer your encoding scheme.
- dustinmoris 7y ago> Instagram uses basic base62 with a very standard alphabet > ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789-_ This is base64. Base62 would be without the "-_" (26 upper case letters + 26 lower case letters + 10 digits = 62, if you include two more characters then it's base64, regardless if the characters are url-safe like -_ or /+ like in the standard base64 encoding). EDIT: Also worth noting that with things like base64/base62/etc. you can either encode a byte array into a string of known characters, or if you deal with integer values only then you can compress it much more by converting from a base10 number to a base62 (or similar) number. For example, "1000" base64 encoded yields "MTAwMA==", but if you'd convert the base10 number 1000 into a base62 number then you'd get "G8". It's like converting between base2 and base10 (binary to dec) or base16 (hex) to base10 (dec).
- doh 7y agoMy shameful confession is that I never counted the characters. I was mentally looking for = and + and subtracted them from the number. Thank you for correcting me
- 3xblah 7y ago"This encoding may be referred to as "base64url"." https://tools.ietf.org/rfc/rfc4648.txt https://tools.ietf.org/rfc/rfc4648.txt