4 ms·
Yeah, that's a massive problem with the natural language domain all across machine learning. Unfortunately it's very difficult to track down training data for
by cbutner 5y ago
Yeah, that's a massive problem with the natural language domain all across machine learning.
Unfortunately it's very difficult to track down training data for chess commentary in the first place, let alone trim down biases. For reference, I was able to gather about 1 million samples, but it really needs a billion.
Hopefully through data augmentation and better general intelligence models we can make better progress on bias issues soon, as that's a huge problem when we start trusting AI models too much in life.
- prezjordan 5y agoAppreciate the honesty here. Pretty wild how natural this model feels with 1 million samples.
- cbutner 5y agoSometimes it seems really accurate (like the cherry-picked GIF in the overview docs) and sometimes really off. I think for the most part, it knows more than it lets on, but finding the right sampling methods (or better yet, generalized search) to generate the best comments is a tough problem because it's difficult to evaluate quality. There's some info on the sampling methods here: https://chrisbutner.github.io/ChessCoach/high-level-explanation.html#covet-sampling https://chrisbutner.github.io/ChessCoach/high-level-explanat...
- a_t48 5y agoYou might be able to kludge a fix to tokenize the output and replace he/him/she/her with them/their. It's not as sexy as the engine outputting the correct words, but it should get the job done.
- cbutner 5y agoYes, in this case as long as they still agree when it actually names people, I don't think it would be too difficult. There may be factors I'm not considering though. Harder would be more general models like GPT-2 and GPT-3.
- a_t48 5y agoSingular "they" doesn't care about the gender of the person named, so it should be good.