3 ms·
I know you're making a general point here, but for the example given, I think it's generally best to refresh your dev data from production regularly (maybe afte
by bendtheblock 17y ago
I know you're making a general point here, but for the example given, I think it's generally best to refresh your dev data from production regularly (maybe after each large release). That way you're always testing on data 'close' to the current production state, which would help you avoid any data-related issues like this.
- patio11 17y agoI think that is a good idea if you're able to get away with it. Scribd can. We can't do it at my day job. There are hundreds of different ways to break the personal information privacy law if we screw up anonymization of test data. For example, supposing we just take the naive approach and overwrite all names, addresses, emails, phone numbers, etc etc etc. Should be fine, right? Except, uh oh, the data is still personally identifiable: the data will tell you that a female student who took CS103 and English 101 last year was given a semester of medical leave. If a copy of that dataset leaks and someone in the department realizes "Hey wait, the only person in that double major is Hanako... medical leave... Hanako was pregnant last year?!?", then our company just made the front page news, we made our customer look horrible (and likely owe them and Hanako several tens of millions of yen in we're-so-sorry money), and we just broke the information privacy law something fierce. Incidentally, engineers not treating test data with the same "This CD is nuclear waste" precautions we treat the production data set is a frequent cause of breaches like this. Somebody decides to work from home for the day, gets his laptop stolen, bam front page news. I nearly got in severe trouble for leaving a printout of the student roster on the printer fifteen feet from a door somebody could tailgate through -- the only thing that saved my keister was that I could show that the student roster I printed out was fake. (Lesson learned about producing good test data: don't produce too good test data.)
- bendtheblock 17y agoWhy is this being down voted? The privacy concerns are valid but are the symptom of a different problem, which is controlling machine/network access. Previously I've worked in banking tech and this procedure was always followed (with data scrubbing), but it was also low risk since we were already on a secure network in a secure building. For a startup with sensitive live data though, I can see why it's best to keep prod data out of test environments and therefore lower the chances of laptops being left unlocked or print outs lost that contain sensitive data. It's clearly a cheaper and better solution to keeping everyone in a secure building and limiting their remote access.