3 ms·
This article touches on the more critical and flip side of a significant problem, which I think will hurt many of the incumbents, while whoever figures out how
by bertil 3y ago
This article touches on the more critical and flip side of a significant problem, which I think will hurt many of the incumbents, while whoever figures out how to resolve it will gain a lot of credit.
Facebook doesn’t explain how their ad targeting works at the individual level, but they thought about it and tried—barely. It was perceived as an obstacle to black-box ML, even when there were clear opportunities to show customers (and partners): here’s how the algebra works, here are key dimensions that we have identified explicitly and that the advertisers wanted us to target (age, gender), here are more dimensions and the feedback from showing that ads to other tells us (people who like board games respond more to the ad, and we think you do).
That training, in clear, accessible language that anyone would understand exists: it’s part of onboarding.
The result of not having it is a large chunk of ads are shown to people who literally don’t speak that language (something Antonio GM mentions in his book and that was identified, corrected, broken, and re-identified at least three separate times since). But even benefits would be enormous: it would free what looks like 60% of my ad inventory by not showing gambling ads.
But they went the other way, treating users asking questions that their own employees ask on their first days with suspicion—assuming they knew better than the people they are talking about.
OpenAI is less adversarial, but they must explain how training and testing work with user-relevant examples: “Here’s a conversation you had last week. Here’s the feedback you gave. Here’s how we detected it was relevant to refinement. Here are examples of questions whose answers were changed by your contribution.”
Having a full-text search on their training set might be difficult, but something along those lines could be implemented and reviewed by a neutral third party. Those conversations would lead to insights about where to find more relevant data.
As Simon W. writes, this matters, and you want to respect privacy because it’s an inherent good, but you also want to preserve trust because it’s commercially essential (even though people are quick to forget if they get free stuff). I think you should do it because security by obscurity, or at least the ML equivalent of that, isn’t working.
https://xkcd.com/1838/ https://xkcd.com/1838/ is funny, but the truth is that we can offer some tools for exploration and feature breakdown. Every time I did, we learned so much from those.