4 ms·
> How do you even know all the permutations of personal data that came be stored in the logs. There are possibly infinite possible ways personal information can
by allover 7y ago
> How do you even know all the permutations of personal data that came be stored in the logs. There are possibly infinite possible ways personal information can manifest in logs.
Nonsense. You write the log statements. You know what data structures you are logging.
If you're using some server's built in logging, or some logging library or middleware you don't understand, turn that off until you understand what it's logging.
- bduerst 7y agoGP is referring to pseudonymization, not data structures. Logs that do not contain explicit PII are still rife with pseudo-identifiers that could possibly (but not typically) be used to join activity with PII. For example: - You have one set of logs that stores anonymized click activity - You have another set of logs that stores purchase transactions - Both have millisecond timestamps You could potentially link the click record from one logs database to the purchase record in the transaction database, when during a single millisecond there is only one transaction and one click happening. Now your anonymized click ID and all your click activity is linked to your PII in your transaction. Sometimes it's off by a few millisecond. Sometimes the logs are obfuscated up to the second level, but then you'll still have instances of a single click and transaction in a second. Does this activity still need to be removed from logs, despite not being linked to your PII or even being identifiable? These are the challenges that need to be addressed.
- laughinghan 7y agoYes. Emplify, for example, makes it a point not to reveal averaged responses for subgroups of size less than 5, for similar reasons: https://intercom.help/emplify-insights/en/articles/1731829-confidentiality-policy https://intercom.help/emplify-insights/en/articles/1731829-c... We should be making an effort to take such care with all customer data, even just when storing it. Mistakes are inevitable, of course, so small gaps that are soon fixed should be let off with a warning, with any fines proportionate to the amount of exposure and negligence involved. How would we do that? Maybe have an agency of experts tasked with determining the fines, and allowing companies to appeal those fines in open court. Like what GDPR does. I'm a software engineer who works for a SaaS data analytics startup that has to comply with GDPR. It's not cheap, just like it's not cheap complying with all the laws restricting pollutants emitted by my car, but it's still completely worthwhile. (My employer is not Emplify, although we are a customer of theirs. Good service.)
- bduerst 7y agoThat is not how logs work. Did you reply to the right comment? Exemplify is gating database records from surfacing through their UI. These records still exist in the database and are admin accessible. The act of knowing if clustering the data is too much still requires knowing the data - i.e. the data existing.
- laughinghan 7y agoYes, I replied to the right comment, you're getting hung up on an irrelevant detail. I'm saying that the thought they put into anonymizing the data they surfaced through their UI, that same amount of thought should be put into the data we all store. If the data can't be clustered in a way that preserves anonymity, it should be deleted (after the desired aggregate statistics are computed). Emplify probably isn't required to, and so they probably don't. I'm saying they should be required to.
- bduerst 7y agoIt is not known at the time of storing the data if the data being stored will be denonymizable - especially for logs data. Saying, "Just don't store logs data" is a fundamental misunderstanding of how web development works. This data is crucial for operational uptime, debugging, and running an online business. The scope of the data is so large that there inevitably are factors that can be used for denonymization, which is to GP's point. The reason I asked if you replied to the right comment is because logs data is fundamentally different from database records, which is the working example you gave you gave with exemplify.
- laughinghan 7y agoI'm not fundamentally misunderstanding anything. We both understand the problem quite well, you're just refusing to take on the burden of trying to solve the problem. I didn't suggest not storing any logging data, actually. I suggested deleting it. Old, stale logs are unnecessary for operation uptime or debugging and low-value for usability or security investigation. They also cumulatively presents risks to customers. A gay blogger in Russia who used LiveJournal in 2004 might regret their decision now, even though in 2007 when LiveJournal sold to a Russian company, few reasonable people would have foreseen the country's turn towards homophobia later. If LiveJournal had, for example, replaced all IP addresses with cities in historical, pre-2007 HTTP logs, they would have lost nothing of value to them while their customers would be that much safer. If they had gone so far as aggregated statistics of requests and unique visitors per tuple of (user agent, city, timestamp truncated to 15-minute intervals), and then deleted detailed all HTTP logs older than 90 days as suggested by Maciej Ceglowski [1], can you think of anything of value they would have lost? But of course I'm sure they didn't, because they weren't required to put that much thought into the data they stored. I'm saying we should be required to. [1]: https://idlewords.com/talks/haunted_by_data.htm https://idlewords.com/talks/haunted_by_data.htm