12 ms·
Also, how hard is it really to get a sense for the quality of a search and the irrelevance of its results for the vast majority of people? Given how much demogr
by andkon 9y ago
Also, how hard is it really to get a sense for the quality of a search and the irrelevance of its results for the vast majority of people? Given how much demographic data Google pulls in from your history, it seems to be not super hard to notice that actual black girls are repulsed by what was shown to them. It also seems easy to post a notice saying: "Hey, searchers like you have found this query to contain many irrelevant or inappropriate results. Beware!"
- amelius 9y agoTaking it a step further, Google could apply this rule: if users really wanted to see (e.g.) porn, then adding that word to the query would show it; so let's keep the results safe by default.
- nickpsecurity 9y agoI like the simplicity of your idea. However, there's articles on porn that are non-pornographic in nature. Usually people discussing or discouraging it rather than using it. To account for that, maybe make it something like "showporn" or a dedicated search page for it. Then, people Googling for information on porn that's not pornographic still wont get hit with it. Whereas, people looking for it using dedicated search page will exclusively see it. This isnt that different from search engines for specialized databases. I used to use meta-search engines that woukd have checkboxes or selection menus to focus search on specific topics or collections. It was very useful. I still do it for technical papers using the "site:" operator to limit search to known-good sites for them.
- PeterisP 9y agoIsn't the whole thing presumably caused by the fact that out of people using that query, most do want to see naked girls? So what you're proposing is to set the comparably less frequent, minority option as the default one, which is quite strange - generally you'd set the defaults so that usually they would not have to be changed. In general, there's the "safe search" option that's on by default; but if someone has turned it off, then it's reasonable to assume that the default, most common intent of looking for "girls" actually is sexual, and if they were looking for something else (say, "black girls support group") then that would be a comparably rare situation where extra words should be added to the query.
- politician 9y agoThe problem statement is given the sentiment space of reactions to similar clusters of search queries, minimize negative sentiment across the most clusters given resource constraints (time, cost, CPU). This is a hard problem. There are many search queries, there are many clusters of similar search queries that address partially overlapping (topic,audience segment) pairs. Deriving sentiment scores from these segments is difficult without any feedback mechanisms. Choosing which segments to focus on is difficult or counterintuitive and depends on the topic and audience and characteristics of the audience. Once a topic and audience segment is identified for remediation, the task of effecting change is again topic and audience specific (resistant to automation). This is a hard problem. Neither a brute force "human review squad" approach nor an automated "deep learning" approach will provide 100% coverage through all time. Does that mean that they shouldn't try? No, of course not, and they almost certainly try every day. However, it's inevitable that they will be subject to sniping articles no matter what they do.
- andkon 9y agoI think I agree with you overall — there's no way to perform this task perfectly for all time in such a way that it makes this problem disappear. But I want to push back against the specifics of your argument, namely about the cost and resources involved, as well as the lack of feedback mechanisms. I do that below; I wrote those paragraphs before really giving credit to how hard it would be to comprehensively score audience + topic pairs. That sounds like... exponentially growing complexity. So that approach sounds like a non-starter; maybe an alternative is to look at quality metrics on clusters of search queries as a whole, before diving into seeing if there are problems with individual audiences and those queries. Who knows. There are many search queries to watch, but 'black girls' was persistently problematic for a long time. Noticing when a cluster of queries gets that status might be extremely resource intensive if that status was highly ephemeral; it's not. It's work that can be done asynchronously, and run daily at most. Is this simple? Nothing involving software is, imo. But it's probably much simpler than many other machine learning-driven pieces of Google's product. Also, there are tons of feedback mechanisms available and actively used by Google. Every user interaction with search results is available to Google; a lot of these stand in for quality of result: did the person refine their query after seeing crappy results? Did they click a link and then press 'back' really quick? All of these factors already feature in Google's algorithm.
- geofft 9y ago> It also seems easy to post a notice saying: "Hey, searchers like you have found this query to contain many irrelevant or inappropriate results. Beware!" They did exactly this in 2004, when searches for "jew" brought up anti-Semitic website "Jew Watch" as the top result (because, as a linguistic quirk, "Jew" tends to be used in a slur-like way and Jewish organizations tend to use "Jewish," and the algorithm at the time considered those as different words): they placed an ad at the top of the page with an explanation of the "Offensive Search Results," distancing themselves from the content and explaining the "Jew"/"Jewish" thing. https://en.wikipedia.org/wiki/Jew_Watch#Google_Search_results https://en.wikipedia.org/wiki/Jew_Watch#Google_Search_result... http://sethf.com/anticensorware/google/jew-watch/jew-watch-com-yes.gif http://sethf.com/anticensorware/google/jew-watch/jew-watch-c... (and yes, that's Netscape Navigator!) http://web.archive.org/web/20050123081919/http://www.google.com/explanation.html http://web.archive.org/web/20050123081919/http://www.google....