4 ms·
Hey all, I'm the one of the authors of the conference paper discussed here and was quoted in this. Glad to see it's interesting to HN! Wanted to briefly highli
by nickvincent 6y ago
Hey all, I'm the one of the authors of the conference paper discussed here and was quoted in this. Glad to see it's interesting to HN!
Wanted to briefly highlight a couple points that I think will be interesting to the HN audience.
One of the major goals of the paper is to describe a framework of three "data levers" (ways a group of people can hurt or harm a data-dependent technology). Data poisoning (well known to ML people for a long time) is one of the three "levers". The other two are "data strikes" (withhold future data and/or delete past data via deletion request) and "conscious data contribution" (ala conscious consumerism — give data to a firm you support and want to compete with incumbents).
A major point in the paper is that there are some big differences in terms of barrier to entry, legal considerations, ethical considerations, and ability for a data lever to be impactful. Basically, for any given company + technology, there's probably a particular data lever that's a "best fit". It might hard to organize a large enough "data strike" that will meaningful hurt a huge company's search engine, but conscious data contribution could help improve a competitor (esp. if that competitor focuses on search verticals). On the other hand, data strikes could be really great vs. facial recognition, because there's precedent of forcing companies to delete actual model weights ([https://www.theverge.com/2021/1/11/22225171/ftc-facial-recognition-ever-settled-paravision-privacy-photos](https://www.theverge.com/2021/1/11/22225171/ftc-facial-recognition-ever-settled-paravision-privacy-photos) https://www.theverge.com/2021/1/11/22225171/ftc-facial-recog...).
Another point is that there's some nice connections between levers. On the topic of data poisoning defenses: if you've been feeding poisoned data, and get caught (quite likely for naive attacks, as noted below), the company deletes your poison and you've just been "reduced to a data strike".
A final point: the paper discusses implications for folks who work in ML, design, HCI, and policy. There's great opportunities to build to tools to support data leverage, and for ML researchers to "bake in" data leverage (e.g. compute a performance v. dataset size learning curve to characterize how "vulnerable" a system is to data strikes). Also, there's huge potential for win-wins with privacy regulation: data deletion and data portability both enhance the public's leverage.
I'll end this long comment now, curious to see what others think (and appreciate all the comments already here!)
- nickvincent 6y agoShould also add, here's the pre-print link for the full paper: https://arxiv.org/abs/2012.09995 https://arxiv.org/abs/2012.09995