6 ms·
Generally speaking, "the algorithm being a mystery" seems to be causing a lot of societal ills. What news we see on Facebook, what sources we see on Google, wha
by jsonne 7y ago
Generally speaking, "the algorithm being a mystery" seems to be causing a lot of societal ills. What news we see on Facebook, what sources we see on Google, what pressure content creators feel. All of it causes a lot of anxiety and really only profits a small handful of companies by allowing them to dodge moderation duties. Moving companies that rely on UGC closer to a publisher probably is a good thing for the world (and content creators).
- eloff 7y agoSerious question, are we at the point where some algorithms are so important in our society that they should be a public good? Should the government require that Google search, Facebook and YouTube recommendation algorithms be published? And should changes to those algorithms need to meet certain criteria at by regulators, in order to prevent ill effects on society? I'm thinking there must come a time when our machine algorithms are so impactful in society that we can't afford to leave them in the stewardship of a single private company with only profit motives. I think we probably reached this point a while ago.
- eachro 7y agoI'm not entirely sure how helpful it would be to see these algorithms since many of them rely on machine learning models underneath the hood (IE: a model that predicts the likelihood a user will click a given video). Is it useful for the average person to know that some gradient boosted tree/deep learning model spat out a probability estimate? More information certainly does not hurt, but I doubt the average citizen/gov regulator can do much with this knowledge. These are not the kind of static bfs/dfs/quicksort/etc algorithms we learned in undergrad that can be dissected so easily. The ML models that underlie the recommendations/ranking "algorithms" are constantly changing based on the data they're trained on. Does this constitute an algorithm change? Disclaimer: this is just my understanding of how these newsfeed type algorithms work based on conversations with friends who work on these teams at FB/GOOG. Please do correct me if I got something wrong here.
- eloff 7y agoYeah ML changes the game here, not even the company really understands the machine singing the music to which we're dancing in that case. It's a strange situation.
- jasode 7y ago>Should the government require that Google search, Facebook and YouTube recommendation algorithms be published? Many folks often wish that ranking/recommendation algorithms like PageRank search engine, or Youtube "watch next", or email spam heuristics, etc, be revealed to the public. E.g. the exact algorithm is required by law to be published on Github. That would hurt society more than help because once the algorithms are public, the code's inner workings can then be gamed and abused. That's the paradox or contradiction: for the algorithms to be effective, they have to be kept a secret to prevent gaming. Recommendation/ranking algorithms are not to be confused with "security by obscurity" memes. Yes, you can publish the exact algorithm to derive a SHA-256 but that's not the same domain as a ranking algorithm. For Google to transition the algorithms from basic PageRank to Panda to Penguin to Hummingbird, etc ... the inner workings must be kept secret to be effective. I know it drives people crazy that the algorithms are secret but you can't make them public in a way that stops abusers from using that knowledge against us.
- dsfyu404ed 7y agoI don't see why the gaming problem you describe matters in anything but the absolute shortest term. If algorithms are gamed they'll have to be modified to be resistant to gaming in order to deliver effective results. As the game of cat and mouse goes on it would result in us refining our algorithms to the point where they actually measure what we want rather than an approximation of what we want. Eventually "gaming the algorithms" would be the same as producing high quality content. Yes refining algorithms to that point would be hard work but it would get figured out.
- jasode 7y ago>If algorithms are gamed they'll have to be modified to be resistant to gaming As far as I know, nobody in computer science has come up with a ranking algorithm that can't be gamed to render the algorithm useless. (Unless the so-called algorithm is a trivial case such as the runner with the fastest time across the finish line is the "winner" of the race.) For a multi-criteria algorithm that approximates a hazy idea like "video quality" that uses various signals as proxies for quality, the algorithm's exact source code must be kept a secret. Yes, you can publicize broad outlines of what the algorithm does but you can't reveal the actual code (or the actual machine learning weights, etc). Yes, an algorithm like a SHA256 hash source code can be made transparently public and still can't be "abused" such that a bad actor can make the string "DEADBEEF" appear anywhere they want in the output hash. This class of algorithm behaves differently than ranking/recommendation algorithms.
- aaronax 7y agoInteresting idea, and reminded me of how I sometimes think about the effect a ban on unnecessary "user eyeball time optimization" would have. We have a bunch of companies that run A/B tests all the time, and of course either A or B result in more time spent watching or less subscriber cancellations. So then they keep that change. Hence you get a product that over time is literally designed to be as addicting as possible...that can't be a net positive for society in my mind. An interesting idea I have (not realistic of course) is that companies would have to "send their customers away." Like the mountain lion in that Trailer Park Boys episode where "if you love something you have to set it free," a positive way to service your customers is to help them get done with what they want to do as fast as possible. "But my customers want to be entertained for X hours." Okay, provide a UI for them to find stuff to watch but do not autoplay videos (for example).
- snowwrestler 7y agoI think there's a fundamental misunderstanding about how these systems work in 2019. There's not "an algorithm" that can be published and inspected, like good ol' Pagerank. It's probably more accurate to think of it like thousands of slightly different algorithms knotted together, tied to an enormous database that selects with algorithm gets used for which person at which point in time. The knot and the database are constantly changing. You could perhaps publish the machine learning algorithms that created the recommendation engine, but those won't tell you anything useful without the training data (which, for these systems, is absolutely huge and complex). I'm convinced that a big reason these social platforms haven't fixed their recommendation engines is that they actually have no idea why they recommend the things they do. The systems are trained on "engagement", without regard to whether it's good or bad, happy or sad, constructive or destructive. It's just: which thing collects the most clicks from which people. In a way, this makes these recommendation engines mirrors into human nature--which sounds cool, but is actually boring and banal. We already knew that it is easier to make someone unhappy than to make them happy. And we typically want people to build things that deliver happiness--even though we know it's harder.
- xamuel 7y agoYou think it's bad now, just wait til they start using ML to automate the peer review process at scientific journals. Dark days ahead!
- brandmeyer 7y agoI think there is a real opportunity for an editorial ML system to help enforce a journal's linguistic style. I think you could train ML to recognize the difference between active and passive voice, for example. But using ML for controlling content? Please. I very much doubt anything like that is going to happen.
- Tarsul 7y agoand if you spin the argument further, that also describes why AI is overvalued (at least in terms of how hyped it is in reporting): If you have no chance of knowing how the result (e.g. of an algorithm) came to be, how can you trust it? In the long run, you can't.
- wesammikhail 7y agoshort answer is: you cant. You don´t even have a way of benchmarking it against deterministic models. All this AI heuristics stuff, while useful in some domains, have become more dystopian than utopian.