3 ms·
Bias of various forms in the datasets we use can absolutely be a big issue, and this is a pretty good summary of some of those areas. However, I think it's impo
by armoredkitten 6y ago
Bias of various forms in the datasets we use can absolutely be a big issue, and this is a pretty good summary of some of those areas. However, I think it's important to look beyond just the data and also look into the assumptions and choices we make regarding models, performance metrics, etc.
I came across a good Twitter thread[1] explaining some of these other types of bias -- a lot of them come down to various ways in which model decisions end up impacting performance on the "long tail" of data (i.e., the less frequent categories and groups) long before they impact the bulk of the distribution. This means overall performance may be minimally impacted (or even improved), but performance for subgroups can be drastically reduced.
Anyway, the thread is definitely worth a read, and it links to many sources for further reading.
[1] https://twitter.com/sarahookr/status/1361373527861915648 https://twitter.com/sarahookr/status/1361373527861915648
- beluis3d 6y agoCompletely agree. There are three primary forms of bias: Human Bias, Data Bias, and Algorithmic Bias. A better solution is to improve existing machine learning models with existing solutions (e.g., domain adaptation, domain generalization, discovering latent domains). This will improve the overall performance.