4 ms·
You don't have to be a machine learning expert to understand that no classifier is going to be correct 100% of the time. The laws against divulging PII don't co
by staticautomatic 7y ago
You don't have to be a machine learning expert to understand that no classifier is going to be correct 100% of the time. The laws against divulging PII don't contain exceptions for classifiers goofing.
- derision 7y agoAnd you don't have to present criticism by calling a group of people you don't know idiots.
- thruflo 7y agoThat’s not how it works. Synthetic data is entirely artificial rather than transformed, so you’re not worried about “missing some PII”. See for example the videos on https://hazy.com/product https://hazy.com/product Disclosure: Hazy cofounder.
- staticautomatic 7y agoWhy would you need "to anonymize vast amounts of data — so that it’s no longer tied to customer information" or "appl[y] access policies" if the data contain no PII? Presumably the ML is anonymizing the data and the access policies are necessary because the data contain PII.
- thruflo 7y agoYup, people tend to confuse concepts and refer to synthetic data as anonymised data. They are very different things. Anonymised data or redacted data are transformations of a data set that _hopes_ not to leak too much PII / sensitive data. People don’t use ML to anonymise but they do use ML to classify as a first step before splatting or generalising. In that case, its absolutely right that the ML classifier not being 100% results in PII leaking. This is a key reason why anonymisation and redaction are widely seen as problematic and are being replaced by synthetic data and, maybe in future, homomorphic encryption.
- thom 7y agoOut of interest, can Hazy learn temporal relationships in streams of data?
- thruflo 7y agoYup. We focus on financial services, so time series data is a key requirement / capability. There are nuances — like static vs dynamic distributions and whether the downstream model is actually just aggregating anyway. I’m my username at hazy.com if you’d like to connect.
- woeirua 7y agoSynthetic data uses the distribution(s) of the underlying dataset(s) to generate a totally new dataset that in theory has the same statistical properties of the original dataset. The rub is in making sure that you're actually synthesizing a valid statistical representation of the original dataset, including joint distributions. Otherwise, you wind up with models that won't generalize back to the original data.