5 ms·
Some problematic scenarios: - How do you identify what is customer data? There may be information stored in logs somewhere. Do you now have to write log pars
by bcheung 7y ago
Some problematic scenarios:
- How do you identify what is customer data? There may be information stored in logs somewhere. Do you now have to write log parsers to extract personal data for everything that previously you just stored for general debugging and security purposes? How do you even know all the permutations of personal data that came be stored in the logs. There are possibly infinite possible ways personal information can manifest in logs. How do you ensure compliance with something when you don't fully understand what can come out of it? Any engineers now must fully understand the consequences of anything they log and design delete mechanisms for it. This extends to any 3rd party software you use that generates logs. You must now fully and deterministically understand your entire system just to comply with this law. Such a request is essentially NP-complete.
- How do you prune said data from logs?
- How do you delete data that are archived in write only media formats and/or that are in cold storage somewhere? You'd have to physically destroy the media and make a copy of everything minus the part you want to exclude. This dramatically increases archive storage complexity and cost.
- darknoon 7y agoYes, if you log personal information like IP addresses you need to have a plan to delete it. Maybe storing it long term is a liability not a benefit. Don't bring complexity theory into it.
- codeulike 7y agoIn Europe, we have GDPR which is broadly similar. And the answer to your question about logs is basically "Tough shit. Personal data is important and if you've been leaving it in logs all over the place then you're going to have to sort that shit out". The Backups question is a bit more complex. One source I've seen: "According to France’s GDPR supervisory authority, CNIL, organisations don’t have to delete backups when complying with the right to erasure. Nonetheless, they must clearly explain to the data subject that backups will be kept for a specified length of time (outlined in your retention policy)." Paired with that is that if you're keeping data (or backups) for any length of time beyond the immediate needs of the customer then you need to be able to justify it.
- allover 7y ago> How do you even know all the permutations of personal data that came be stored in the logs. There are possibly infinite possible ways personal information can manifest in logs. Nonsense. You write the log statements. You know what data structures you are logging. If you're using some server's built in logging, or some logging library or middleware you don't understand, turn that off until you understand what it's logging.
- bduerst 7y agoGP is referring to pseudonymization, not data structures. Logs that do not contain explicit PII are still rife with pseudo-identifiers that could possibly (but not typically) be used to join activity with PII. For example: - You have one set of logs that stores anonymized click activity - You have another set of logs that stores purchase transactions - Both have millisecond timestamps You could potentially link the click record from one logs database to the purchase record in the transaction database, when during a single millisecond there is only one transaction and one click happening. Now your anonymized click ID and all your click activity is linked to your PII in your transaction. Sometimes it's off by a few millisecond. Sometimes the logs are obfuscated up to the second level, but then you'll still have instances of a single click and transaction in a second. Does this activity still need to be removed from logs, despite not being linked to your PII or even being identifiable? These are the challenges that need to be addressed.
- laughinghan 7y agoYes. Emplify, for example, makes it a point not to reveal averaged responses for subgroups of size less than 5, for similar reasons: https://intercom.help/emplify-insights/en/articles/1731829-confidentiality-policy https://intercom.help/emplify-insights/en/articles/1731829-c... We should be making an effort to take such care with all customer data, even just when storing it. Mistakes are inevitable, of course, so small gaps that are soon fixed should be let off with a warning, with any fines proportionate to the amount of exposure and negligence involved. How would we do that? Maybe have an agency of experts tasked with determining the fines, and allowing companies to appeal those fines in open court. Like what GDPR does. I'm a software engineer who works for a SaaS data analytics startup that has to comply with GDPR. It's not cheap, just like it's not cheap complying with all the laws restricting pollutants emitted by my car, but it's still completely worthwhile. (My employer is not Emplify, although we are a customer of theirs. Good service.)
- deleted 7y ago[deleted]
- ska 7y agoThere are a (very) few hard edge cases. But most of this stuff is easily addressed by 1) having good designs around data handling in your system and 2) treating customer data and PII as data you only have limited rights to in the first place.
- platz 7y ago> engineers now must fully understand the consequences of anything they log I don't think this is quite the dichotomy you make it out to be. So we can create optimizing compilers, but we can't figure out what to log? This seems like a problem of never having motivation to solve the problem before. "We can't do that, it's too hard" is often a mea culpa I'm industry when they oppose regulation. Then they will come up with a solution from having actually spent some effort to actually think of potential solutions.
- Nasrudith 7y agoOh we can figure out what to log - and figure out how it would suck donkey balls. Just like the other "too hard" areas like the magic golden key backdoor. It can trivially backfire to make things less secure if it is poorly defined which is generally a given. The GDPR made exfiltration easy as an account compromise - one could argue it is an acceptable trade off for transparency but the regulators must bear full responsibility for their constraints. "We can't log sensitive customer data slowing down debugging and worsening data integrity" is one thing but now imagine "can't log customer IDs in a read only way as sensitive information" oops there goes a lot of useful auditing information as it is excluded or forced to be writeable.
- platz 7y agoIt's to hard to do anything so lets do nothing.
- JackRabbitSlim 7y agoALL of those scenarios are only problematic because the design of these megalithic services didn't even consider things like the basic privacy of their livestock/user base. Think this law is going to allow me to require 7-11 to delete me from their DVR records? Not part of the business model, it's a matter of security. Nice straw man though. Their terrible system design can't handle "FROM PornPrefs DELETE SSN,Name,Address WHERE SSN LIKE "999-11-2222"". They brought this on themselves.
- laughinghan 7y agoDo you now have to write log parsers to extract personal data for everything that previously you just stored for general debugging and security purposes? ... Any engineers now must fully understand the consequences of anything they log and design delete mechanisms for it. YES YES YES YES YES. Are you not already doing this for passwords, credit card numbers, and social security numbers? Such a request is essentially NP-complete. I think you mean undecidable, or equivalent to the halting problem, or subject to Rice's theorem. NP-completeness is irrelevant. I think you'll find that HN is the last place you'll win arguments by inaccurately using technical terms in the hopes that it will go over other people's heads, Legally Blonde-style (https://www.youtube.com/watch?v=8rNVaY7Stt4 https://www.youtube.com/watch?v=8rNVaY7Stt4). To anyone who knows what they're talking about, this an obviously nonsensical argument. It's similarly undecidable to verify whether the data that you expose publicly contains customer data, or customer passwords, or your own passwords, but you do it anyway, by restricting your engineers to only write and deploy code that they understand.