2 ms·
We're building a consent framework API so our customers can consent for personal data use. Data is then cleaned and transformed (ETL) from personally identifiab
by unkoman 9y ago
We're building a consent framework API so our customers can consent for personal data use. Data is then cleaned and transformed (ETL) from personally identifiable to pseudo-anonymised. The data is also separated into two separate encrypted storages for anonymous and pseudo-anonymised data for generalisation and separation. Random (important) hashed identifiers are created and put into a metadata service which is used as a lookup-table. If right to be forgotten is invoked, the data is disassociated from the pseudo-anonymised and personally identifiable data thus making it anonymous.
Important is also how you handle data analytics and this is why we're deploying high restrictions on raw data. Analytics will only be able to be done through an analytics service which can give the employees access to only certain parts of the data which is approved for the use-case. We're using Apache Sentry for fine grained role based authorisation to data and metadata and a directory services for user auth.
Things we've learned:
* Minimise data usage
* Don't use personally identifiable data
* You will need to be able to prove consent when it comes to data usage and it cannot be consent by default, it has to be opt-in
* Log all data access so that use cases can be proved. This needs to be evaluated and audited
* Encrypt in transit and at rest
* Centralise mapping for all data