3 ms·
That's quite amazing! One thing is that the data center ids are in the tweet ids, so it could be used to get a rough location of Twitter users.
by aboutruby 8y ago
That's quite amazing!
One thing is that the data center ids are in the tweet ids, so it could be used to get a rough location of Twitter users.
- aboutruby 8y agoI see there is interest in this observation so I used a little ruby (.to_s(2)[-22..-17].to_i(2)) to get the datacenter id. Then ran it on a few Twitter accounts: https://pastebin.com/w8Dnj5kM https://pastebin.com/w8Dnj5kM It does work, and it's going to be hard to patch edit: I realized you don't only get one location but the whole location history of a Twitter user. Also locating Twitter's data centers as it doesn't seem to be public information
- stuck_in_matrix 8y agoThis is really interesting. When I did the original analysis on datacenter / server ids, I didn't think about correlation with user accounts. Nice observation!
- Sommer 8y agoShould be pretty simple to check the tweet geo or location mentions per server id to see of they imply a correlation with geographic area of the server. Then you're just one hop to knowing where (or where not) other tweeters are.
- aboutruby 8y agoI was curious and just did that :) https://gist.github.com/localhostdotdev/48ed13972c3e5391a47f8e3dd7b9e0dd https://gist.github.com/localhostdotdev/48ed13972c3e5391a47f... (small sample of ~1500 localized tweets)
- jaytaylor 8y agoThey could "fix" it by periodically rotating the DC and server IDs. I wish it was not so easily fixable, because this will break the key space reduction trick ;), but unfortunately such a solution is feasible and would come with the side-effect of drastically increasing the required scanning space.
- atian 8y agoWhat are you guys even smoking. The ID segment is 5 bits long. You have an extra server ID bit making the datacenter results more significant than they are.