6 ms·
How does that work when it comes to private information like emails, search histories, etc?
by joekrill 4y ago
How does that work when it comes to private information like emails, search histories, etc?
- lazide 4y agoIt’s been a decade since open access like that was a thing at Google. Some high profile incidents made the company lock it down. It’s probably one of the most locked down environments now outside of a SCIF.
- deleted 4y ago[deleted]
- zaphar 4y agoYou encrypt or mask, usually mask, that data out of the data lake.
- jrockway 4y agoThe standard technique, at least when I was there, is to encrypt with a per-user public key whose decryption key is only present when your application has the user's login session. Basically, your application can't decrypt the data unless the user is making a request to your app. (There are a lot of potential attacks here, like modifying the code to exfiltrate data when the user is present. Hopefully those are easier to detect than just putting the data in a database unencrypted, though.) I worked on Google Fiber. We sent all the system logs from our WiFi routers (etc.) back to us, mostly to get to the bottom of the most obscure WiFi bugs. This could have been a privacy disaster, as Linux logs all sorts of PII like the MAC addresses of devices that are connected to the router. My team modified the Linux kernel to not do this, and added filtering before upload to XX anything that looked like a MAC address. That way, when some PM came up with a horrifyingly bad idea like "hey, let's make a social network based on who connected to your WiFi", that data didn't exist. (This predated widespread implementation of MAC address randomization. We took care with your PII, but nobody else did, so phone vendors simply stopped giving access points PII. Probably the right solution, honestly.) This wasn't just us that did things like this, it was institutionalized. There was a privacy team that had to review and sign off on this sort of thing. If you wanted to run something in production, or retain any sort of data, someone that actually cared about user privacy would be taking a look. HN kind of undervalues how much people actually cared; there is some trade off between usability and security, and nobody is going to use a product that can't be used, and no corporation can really stand up to the government the way they want. I thought it was a decent balance, though. Obviously, some things can't be redacted like we did, and you should be careful about who you share those things with. Location history comes to mind as the most obviously abused feature. I thought it was super cool to have "you visited this cafe 2 years ago" on Google Maps, but it comes at the cost of fishy warrants for "anyone within 1 mile of this minor property crime". I'd rather not be on those lists, so I don't use it. (But be careful, your cell phone provider surely has this information, even if the Vice article is about Google. They just don't use it for anything useful; it's ONLY a dragnet.)
- delroth 4y agoThis was pre-GMail, which was the first big Google service storing personal user data.
- dekhn 4y agoregular users of GFS (developers) couldn't trivially access that data, but certainly up to a point it was possible for expert users to troll that data (if you were caught, you'd be fired). Eventually better protections were applied. I worked with a guy who had built part of search and later made the gmail "lawyer search" feature (legal discovery). He said that any interest in looking at people's emails went away after having to spend hours going over various emails involved in court cases (typically done in a room with several lawyers) to make sure the ranking algorithms could surface emails demonstrating illegal behavior
- DannyBee 4y agoThat sort of stuff was not really stored in GFS directly, it is stored in storage systems that build on top of GFS. That has ~always been the case. So while GFS might have had, say, open access to the map data and tiles, or the crawled web pages, user data isn't accessed that way (and this was standardized and enforced very strongly, even 16 years ago when i started).
- syscomet 4y agoDefense in depth. First, the common RPC identity and authorization system provided RPC identity to the GFS API calls, and there was a common user/group system. GFS had ACLs and so despite what was said up-thread most people could _not_ access the data in the Gmail GFS cells. The team membership vs runtime user setup also ensured that Gmail SREs using their own credentials could not directly access the GFS cells (but once you have root on the storage boxes, that somewhat disappears). Second, on top of that you have encryption, using an encryption key service which would spot anomalies in things asking for the decryption keys to decode the stored data.
- jeffbee 4y agoThey publish a bunch of whitepapers on this stuff, including how storage encryption keys are unwrapped on behalf of services: https://cloud.google.com/docs/security/encryption/default-encryption https://cloud.google.com/docs/security/encryption/default-en... How services authenticate each other: https://cloud.google.com/docs/security/encryption-in-transit/application-layer-transport-security https://cloud.google.com/docs/security/encryption-in-transit... And how insider risk is mitigated by monitoring the provenance of production software: https://cloud.google.com/docs/security/binary-authorization-for-borg https://cloud.google.com/docs/security/binary-authorization-...
- UncleMeat 4y agoThe insider risk stuff always was really cool to me and, IMO, represents ways that the big tech companies do way more than everybody else in this space. BCID can be a huge pain in the ass but being able to say "hey, it actually would be pretty tricky for a single disgruntled employee to execute code to steal user data" is quite powerful.