2 ms·
> It is the dataset used for training that is often at fault, not the source code of the model This remark is absolutely spot on. This is why I don't like the
by datastoat 5y ago
> It is the dataset used for training that is often at fault, not the source code of the model
This remark is absolutely spot on. This is why I don't like the language around 'algorithmic bias' and 'algorithmic accountability' -- it puts too much attention on The Algorithm, which lets big tech companies deflect us from scrutiny of their training data.
Your comment about the conflict with privacy is also spot on. Here's a paper (shameless plug!) about it: Show Us the Data: Privacy, Explainability, and Why the Law Can't Have Both [1].
What's interesting about the GDPR is that it's not purely about what regulators do: it's legislation that grants rights to people, rights that can be litigated in court. If I as a GDPR data subject demand an explanation of an ML decision about me, as I'm entitled to do, my lawyer could argue "the source code isn't an adequate explanation, we want to see the training data too". Then it'd be up to the court to make a ruling about my lawyer's request, not up to the regulator. The paper [1], co-authored by a data scientist (myself) and a lawyer, thinks through the privacy / explainability conflict as it might play out in court.
[1] https://www.gwlr.org/show-us-the-data/ https://www.gwlr.org/show-us-the-data/