3 ms·
The problem is you cant feed the ML algorithm training data based on what your company currently looks like, you have to feed it an idealized set of what you wa
by noetic_techy 8y ago
The problem is you cant feed the ML algorithm training data based on what your company currently looks like, you have to feed it an idealized set of what you want it to look like. It almost needs to be fictitious training data to hide the ugly bias that's already built in.
I don't think this will ever work. There is too much variability in resume wording that correlates to gender and even culture of origin even when you take out names and any other protected class identifying markers. The Dutch tried this and ended up with less diversity.
I'm going to go out on a limb and say you almost want to leave all that identifying data in, but put each candidate into buckets with separate rating algorithms trained against only that "type" of candidate. The top candidates from each culture, and the top candidates from each gender, etc etc, however you want to do it. Feed them into a picking algorithm that builds a composite of what you want your team to look like diversity wise based on the top candidates from each bucket, and go from there.
Don't take my opinion seriously, I'm not an ML guy.