8 ms·
Should anyone have access to the actual data at their company? I feel like this is an indication of maybe scrubbing the said data is a "must" before it goes int
by gtoprak 7y ago
Should anyone have access to the actual data at their company? I feel like this is an indication of maybe scrubbing the said data is a "must" before it goes into the hands of employees.
Then again, what type of side-affects would that have on the quality of the products moving forward.
- fastball 7y agoSomeone needs to have access to the data...
- pfooti 7y agothat's not entirely true - it's possible for all the data of that nature to be in an audit-logged data store, and to require specific business cases for access to the data. So you can only see a user's private information if there's a specific bug you're working on, and even then the access is audited. I mean, the data is still accessible, it's just not easy to get it willy-nilly without setting off alarms.
- throwawaybbb 7y agoSomeone needs to build and maintain the audit log software.
- JoshTriplett 7y agoTrue, but you don't need access to production data to build an audit logger. And any commits to it should involve code review.
- justinclift 7y agoBackup/restore of audit data can make things complicated. eg the ability to overwrite incorrect/damaged audit data (etc) So far, I've not yet come across a system where some level of direct admin access isn't needed for at least "last resort" situations. (obviously, only available to a very specific set of trusted people)
- save_ferris 7y agoSomeone has to scrub the data, though... Agreed, data governance is important for any company to get right, but someone has to have DB access in order to manage it. When you factor in that so many companies derive revenue from the data they generate, then it gets harder. If you’re Twitter, how can you build the services you need as an architect or data scientist without the actual data?
- CydeWeys 7y agoThe scrubbers do not need raw DB access. All they need is a strictly audited web UI. Very, very few people need root access and the ability to see all raw data. For example, at Google, this is a tiny number of senior SREs and that's it. Your average employees should be using audited UIs running through service accounts that have restricted permissions.
- AmericanChopper 7y agoI’ve achieved this without so much hassle at a smallish/medium sized company (~60 engineers). A typical engineer had access to anonymized data, only a few had access to query raw data. All their queries were logged with justifications. They also wouldn’t query raw data as a matter of BaU (outside of some very rare situations), it was pretty much only ever done when making changes to the ETL pipelines. The solution we had wasn’t particularly difficult to set up, and actually made life much easier for everybody, because we’d provided everybody with a very easy to use interface for the data they wanted (much better than the old school shelling into a DB to run your arbitrary SQL statements).
- big_chungus 7y agoAny chance you'd be willing to share more about your setup, at least high-level?
- AmericanChopper 7y agoWe had a data warehouse where we sent pretty much all the data we had in the organisation, and redash in front of it to allow query access, reporting, etc... Everything in the data warehouse was anonymized, and people only had access to the schemas they needed (though this was defined quite broadly). Anonymization was handled by our ETL pipeline. When we first set it up, the requirements were pretty simple and we just wrote a little java app to do it. This scaled pretty poorly, and the team ended up putting a proper ETL product in there. I can’t remember which one they used, but there’s a lot of perfectly decent products in that space (even some open source).
- skywhopper 7y agoWho does the scrubbing if not Twitter employees?
- noident 7y agoYou can and should restrict the number of people with access to the data, but in a tech company there's always going to be a significant number of people with direct access to the raw data. Software engineers on an on-call rotation, full-time site reliability engineers, data analysts, maybe even some external contractors... maybe this isn't the case for Twitter, but many companies also have to put complete trust in their cloud provider or datacenter, which puts even more people in the loop. Even if you follow best practices with access control, in the end you're always going to have a group of people who you need to trust with access to folks' personal data. Maybe the solution is better audit logging and even tighter access, but I'm not sure leaks of this nature are preventable.
- ForHackernews 7y agoYou can significantly reduce these problems by decentralizing data and moving away from giant platforms that make such enticing targets for espionage.
- CydeWeys 7y ago"but in a tech company there's always going to be a significant number of people with direct access to the raw data" This is absolutely not true. This number can and should be reduced to an absolute minimum number of people.
- bmiller2 7y agoThe absolute minimum number of people may still be significant
- tristanstcyr 7y agoSure, but it comes at a cost. Companies rarely push the limits of these kinds of policies because customers are not willing to pay for them.
- rapind 7y agoMaking the penalties for breaches far more severe would be a good place to start though.
- mirimir 7y agoIf they didn't collect PII, there'd be nothing to leak. I mean, we all know that it's insane to store plaintext passwords. So why is it necessary to store anything as plaintext?
- bschwindHN 7y agoYou never want to be able to retrieve a password after hashing and storing it in a DB. You _do_ want to be able to retrieve text content after possibly encrypting and storing it, otherwise it's useless.
- mirimir 7y agoIt's no more useless than a hashed password is. Depending, of course, on how you want to use it. One could arguably build chat apps and social media that retained no PII. In my opinion, providers retain PII primarily in order to monetize it. But the problem is that it becomes toxic waste. And it always leaks, eventually. Putting users at risk, and damaging providers' reputations. Consider the Tox P2P chat app. Each user device runs a Tor onion service. And chats involve only connections among them. Users need disclose no PII. And there's no need for central servers holding PII. Regarding social media, consider all the "dark markets" that have run as Tor onion services. There's no reason why any sort of social media that you want couldn't be implemented similarly. Although there'd be central servers, there'd be no need for them to handle any PII. Indeed, the fact that "dark markets" handle PII is one of their main weaknesses. And it's not even necessary to use Tor. One can achieve substantial privacy and anonymity using nested VPN chains, with far less risk of attracting unwanted attention.
- chrischen 7y agoNo, even US people like police abuse license plate searches for personal reasons. It's not even a matter of nationality. If data is being collected and the only protection is a "policy," then it's being abused or will be.
- noodlesUK 7y agoThere’s a difference between a policy that’s just enforced by an honour system and a policy that is enforced with strong access control, alarm bells, and a paper trail. It’s very possible to encode the sorts of policies that we need to protect peoples data into real restrictions that are strongly enforced. A police officer might need to do some routine queries, but they should probably have gone through vetting about people they are close to and not have access to their data. Additionally, there should be audits performed on a regular basis.
- chrischen 7y agoI think we're arguing semantics here. What I mean by policy is a de jure rule. I think you can have policy while not having any sort of de facto enforcement, and you can have de facto enforcement (encryption), even if you don't have a policy. Policy is never useful because even if there is enforcement, there is never 100% perfect enforcement that beats out cryptographic enforcement, at which point policy is no longer needed. For example Apple can state that your data is end-to-end encrypted and they have no access, and it would be redundant to also have such a policy saying they will not access your data—they can simply say they can't access your data which is a superset of any such policy.
- ChuckMcM 7y agoI don't think it is semantics. There are policies like "You are not allowed to access user data." and there are policies like, "All access, keystrokes, and applications that have access to user data are logged and those logs are tied to employee IDs. Further the logs are audited and there must by a form 505/2 on file for every access that details the need for the access, what was done with the data, and how the data was handled. If the auditors discover an access in the logs associated with your employee ID and there is no matching 505/2 on file, you will be subject to immediate termination and may be liable in civil and criminal court. Your signature below states that you understand these restrictions, you consent to monitoring of your behavior, and will abide by the policies." Strong audit trails, logs that cannot changed by being created in an immutable way, logged access at all terminals and entry points. Combined with a separate auditing group that reports through a different chain of command (like through to the general counsel or something) and you have a policy with teeth.
- abawany 7y agoI liked the access controls at PayPal. Access to data was a function of insider status with real financial consequences (ability to sell stock was restricted much more for higher level insiders), subject to strict controls and auditing, and required need-to-know periodic renewals. I used to think PayPal was fly by night but my work experience there really grew my trust in their access controls and made me much more likely to use them for payment.
- dredmorbius 7y agoThrough contacts I've heard stories of practices at a wide range of establishments over the past 25-30 years. Practices have almost always severely lagged advice, and though specific leading firms or organisations might have strong data hygiene policies and practices, a great many other organisations do not. Through roughly 2000, the principle saving grace was that disk storage was so expensive, and networking so slow, that large quantities of data were unlikely to be found online except in the case of very major organisations. Most financial firms would read data from tape for analysis or marketing programmes, as an example. A major credit card network might have a couple of, say, Sun Starfire class servers onto which a comprehensive union cardholder databset might be assembled and accessed. One friend reported accessing their campus workstation to which a large national medical insurance database was being processed, from the New York Public Library over Telnet (though I believe they didn't actually log in, they did receive the prompt). E-commerce software vendors and systems stored credit card information, which was accessed. Numerous services and datasets fly around all kinds of organisations, with little protection, and were transmitted in unencrypted FTP sessions. Social networks in which NOC addresses were directly accessible from the office network (WiFi access, natch), with millions of members' data directly accessible. There are many ways to get this wrong. Few to get it right. And most organisations lack the staff, capitalisation, or incentives to do the right thing. Google are problably among the best. That leaves open the question of how good they are, and what their past practices have been, even in relatively recent years. Or how they might behave should their advertising monopoly and revenues fail.
- johnpowell 7y agoMy first year at a community college I was doing work study that was part of my FAFSA. I was late getting in there and I had limited options, one was working in the cafeteria, the other was working for the very job placement center that was tasked with finding me a work study job. I got very lucky here. The work study job in the job placement center was turning job listings that were faxed in into html to post on our fresh new website. Nobody at the time new what HTML was. A few years earlier I had bought a "Learn HTML in 24 hours book" and I made a Tony Hawk Pro Skater webpage that listed all the special moves. I used a lot if iframes and thought it was pretty good. CSS wasn't really a thing back then. iframes and tables got the job done. But I got the job and they thought they got very lucky. I worked in this back room with a computer and a fax machine. Jobs listings would be faxed in. I would scan them and let the OCR software try, and then I would clean it up and add some <h> and <b> tags and then do my econ homework for the rest of my shift. But as the digital stuff become more popular they hired another guy to be in the back room with me. Dude was a bit of a creep and a student came in looking for a part time job that would work around her classes. He kept on going on about how hot she was. A few weeks later he was talking about he signed up for a few of the classes the hot girl was taking. Every single student record was available on our computers. Names, address, phone numbers, class schedule, SSN, FAFSA data. It was madness. And I was a lowly fax to html guy.