4 ms·
WOW! This was a really interesting read. I don't even know exactly where to start. I never looked into G+ (and was nowhere near as good as a programmer as I am
by pazvanti 6y ago
WOW! This was a really interesting read. I don't even know exactly where to start. I never looked into G+ (and was nowhere near as good as a programmer as I am now), but those were really cool findings.
Related to UUIDs, it does indeed protect privacy and protects the server from crapping and leaks. It is not enough for security, but it does make it harder. Internally you can still use sequential IDs for the PK so that speed is not impacted.
- dredmorbius 6y agoRight. I'm referring here to the external affordances of UUIDs. Site performance is not my concern ;-) The degree of programming involved in all of these exploits was laughably small. The Ello link literally includes the script I'd run for the G+ data pull, it's a bash one-liner. Parsing mostly consisted of a regex grabbing and isolating the posting date and follower counts from the rendered HTML. The larger point being that if you provide a reference to the UUID set, sampling becomes easy, is ... well, come to think, if you're assigning sequential IDs and I know the upper and lower bounds of that range, I can still sample randomly from within it. The problem of exhaustively searching a large-relative-to-references UUID space also apply, though I'm not sure if/how that might impact internal or site operations.
- saurik 6y agoBoth of you are assuming random IDs protect user privacy. The only thing I can come up with where that to be true (without things that no one was ever seriously suggesting or using, such as a sequential identifier that increments on something per user) is that someone who is unable to see a resource otherwise will know when it was created (and in fact the time when something was created is now information about a thing "at all" when it otherwise might not have been). What am I missing? Otherwise, the arguments against sequential identifiers seem to mostly come down to either people who somehow dislike scraping (which I might even argue is immoral, and I think this great comment from dredmorbiu come at it from the correct mentality; hell, I strongly prefer it--as a potential user--when a site is scrapable) or people who don't want anyone to know "true truths" about their service, with examples such as "our user growth figures" (but these help the site operator, not the users). (I am also surprised to see Parlor brought up even in this great comment about "you didn't really stop me from figuring out what I wanted to know": to the extent to which you want to do "defense in depth"--which I honestly find frustrating as an argument by itself, as you can use that to justify almost anything, including absolutely ridiculous things we tend to claim are "security by obscurity"--you are going to then need to make sure you really really lean into the idea that "we wanted our site to not be scrapable, even to an administrator", which I can't imagine anyone ever actually doing for myriad reasons including how sites usually help search engines.)
- dredmorbius 6y agoFair points, though some counterarguments: - Not all systems make user data (or all user data) visible. UUIDs make guessing or traversing the address space harder. - Certain types of attacks are easier given known, guessable, predictable, or sequential identifiers. UUIDs mitigate these attacks. - Certain types of system confidentiality are generally easier to maintain with UUIDs. (Unless, of course, you then hand over the full listing, as Google did with G+.) - Where merging systems, UUIDs may (or may not) help avoid namespace collisions. (Merging disparate systems with independent UID conventions is a notoriously fraught problem.) On sites helping search engines: so far as I've been able to suss out, Facebook actually leans strongly the other way. I don't seem to get much insight to Facebook by trying to do site:facebook.com limited searches (through DDG, which is to say, Bing, or Google). FB's view seems to be that if you want to see or search content on the site, you create an account and go there through the account to find content. I don't often find myself interested in searching FB myself, but for as large a site as it is, it seems to turn up in results remarkably less often than other notably closed/annoying domains, say, Scribd or Pinterest. Even as of 2015, FB's representation among various traditional and social-media options for a set of arbitrary interest queries was relatively modest (though higher than I'd recalled before going back to look at my results) and I believe (based on impressions rather than focused research) it's become more search-engine opaque since: https://old.reddit.com/r/dredmorbius/comments/3hp41w/tracking_the_conversation_fp_global_100_thinkers/ https://old.reddit.com/r/dredmorbius/comments/3hp41w/trackin...
- ConnorLeet 6y agoPossessing an ID, shouldn’t give access, regardless of whether it’s a numerical PK or a UUID. (Unless that’s a feature, like shareable links) Still need to check if the user should be able to use that ID. If that isn’t implemented, the system isn’t secure, doesn’t matter which path you use.
- dredmorbius 6y agoIf the ID scheme mapping is sufficiently dense, traversal attacks on otherwise obscured namespaces become an option. This might apply to user accounts, posts, payment accounts, or other elements. Security isn't simply about compromising account credentials or access policies. It may be any unintended or unexpected data disclosure, inferred relationships (between accounts, activity, finances, offline attributes, access, reputation, and more), denial of access, stalking or harassment, and more. These might not be unexpected in all cases, but could well be undesired in many instances.