4 ms·
I can't stand the fact that organizations hide behind "algorithms" as an excuse for every unintended side-effect. It's no surprise that algorithms create an in
by Declanomous 8y ago
I can't stand the fact that organizations hide behind "algorithms" as an excuse for every unintended side-effect.
It's no surprise that algorithms create an information bubble for us. I like to use music as an example: the best music recommendations I've ever received have been from friends. I'll describe music I'm listening to and like, and they'll recommend something for me. I'd say 50% of the time I find it unobjectionable, 30% of the time I hate it, and 20% of the time I love it.
Sites have been trying to replicate this functionality using algorithms for a well over a decade. Last.fm used to be a great way for me find music. As far as I could tell, it would just pick music from people who listened to the same artists as you. I'd say I found 30% of the music unobjectionable, 50% of the music terrible, and 20% of the music great. Functionally, this was remarkably similar to my friends recommending me music, because I only care about the music I loved.
Over time their algorithm got 'better' and nearly stopped recommending music I disliked, however I found that it mostly just recommended me music that was in the unobjectionable middle ground of 'meh'. I stopped looking for music on Last.fm because it stopped giving me music I loved.
We've positioned machine learning as a replacement to curation, but what it really does is put forth the least objectionable content. Machine learning lacks the ability to deal with the nuances surrounding controversial or challenging content. It never takes the leap of faith that a person would when recommending you something.
I am really interested to see how this all plays out over the next few years, because I think people are beginning to realize they are addicted to the perpetual stream of 'good enough' content. The network effect will probably protect Facebook for the time being, I assume the first sites affected will be content-driven, like YouTube.
Of course, XKCD addressed the issues I have with algorithms far better than I could ever hope to:
https://xkcd.com/1831/ https://xkcd.com/1831/
- allthenews 8y ago>We've positioned machine learning as a replacement to curation, but what it really does is put forth the least objectionable content. I don't think the problem is with ML. You could easily train a net to make boulder selections with an appropriate data set. I believe what you are seeing is an accommodation for the average person. Most people aren't adventurous, and it makes business sense to get algos working for the middle majority than to cater to tail ends. It also, IMO, makes content at places like youtube less technical and more clickbaity. But again I believe this is somewhat intentional.
- AlexandrB 8y ago>> We've positioned machine learning as a replacement to curation, but what it really does is put forth the least objectionable content. > I don't think the problem is with ML. You could easily train a net to make boulder selections with an appropriate data set. I believe what you are seeing is an accommodation for the average person. Most people aren't adventurous, and it makes business sense to get algos working for the middle majority than to cater to tail ends. You literally just said the same thing the GP did while ostensibly disagreeing with him: ML curated content caters to the lowest common denominator - i.e. the content that most people would be okay with.
- Declanomous 8y ago> boulder selections Would that mean more rock and less bluegrass? (I'm kidding, of course) In all seriousness though, the problem probably is the data set to a certain extent. I am pretty adventurous, true, but at the same time I think the problem is that regardless of what data we collect, we only have the ability to determine if something is inoffensive. All you need to do is look at a site that aggregates reviews like Yelp or Amazon, and read some reviews. 95% of reviews are either 1 star or 5 stars, and it's obvious from reading the text that the rating is meaningless. "I ordered the wrong item by mistake, 1 star" "Food is overpriced, but the bathrooms are clean and wait staff is attractive, 5 stars." I find star ratings most useful because people who rate things 2-4 stars generally have the most nuanced and productive things to say about the product, and actually rate things what they think they should be rated, rather than using the rating system as a vote for "rate this higher/lower." I don't think it's possible to actually come up with good recommendations based on user-reported like/dislike rating. It's not wrong for a user to dislike a song because it reminds them of an ex, but using that as a basis for a recommendation to someone else is entirely useless. Systems like Rotten Tomatoes works really well in this regard, but almost has the opposite problem, which is that it tends to underrate movies with broad appeal, but that's generally not a problem since users will be exposed to those movies anyways.
- tomatotomato37 8y ago>Most people aren't adventurous I would say most people aren't adventurous in everything except their interests, which they are very adventurous in. The problem is that there will always be more people with a meh feeling toward a random subject than people with a genuine intrest in it, and a brute force ML algorithem trained by just throwing unholy amounts of generic personal data at it before being being relased on the giant mashup of random subjects that make up something like youtube won't catch that nuance