3 ms·
Then you would just get something like this [1]. As long as the rewards from spam outweigh the cost of spam, people will try to spam. What I don't understand is
by ceeplusplus 4y ago
Then you would just get something like this [1]. As long as the rewards from spam outweigh the cost of spam, people will try to spam. What I don't understand is that Google solved spam detection 10 years ago, yet Twitter somehow can't solve it.
[1] https://www.youtube.com/watch?v=Y-W0CBOGnnI https://www.youtube.com/watch?v=Y-W0CBOGnnI
- jonathanstrange 4y agoI really wouldn't say that Google has solved spam detection. If you search for any kind of purchasable consumer product, the first page of results mostly contains spam blog entries of very low quality. Most of them are "10 best X" or, even more useless, "50 best X" lists with the advertised product(s) on top of the list. Some of them are clearly auto-generated and still often make it to the top 20 results. That's why many people including me nowadays search for "search term +reddit" when looking for product information. If you include Youtube under the label Google (i.e., Alphabet), it's even worse, there is tons of spam and deliberate political misinformation in the comment sections, although arguably most of these are probably generated by humans. For example, any TV report about Ukraine on Youtube is spammed with pro-Russian troll posts from St. Petersburg and likely hired posters from India. There is also lots of commercial spam in comments. So no, I don't think they've solved the problem. They're using captchas against bots and that's about it.
- mike_hearn 4y agoWeb spam is a separate problem to bot-driven spam posting on social media, despite using the same word. Spam is also not the same thing as posts you disagree with. I went to YouTube, searched for [ukraine war report] and picked the first result that came from TV news to try and check your claim [1]. The top comments are all clearly human, some are claiming to be Ukrainians, and are all anti-Russian. Regardless, your post is a good illustration of the problems with academic bot research. Although researchers claim to be researching spam bots, their methodologies are often defining any post they don't agree with politically as "spam" or from "bots". [1] https://www.youtube.com/watch?v=nx-6Y00MrFo https://www.youtube.com/watch?v=nx-6Y00MrFo
- jonathanstrange 4y agoNo, when I use the word "spam" I do not mean "posts I disagree with", I was referring to posts that are deliberately intended to misinform, are repetitive, or are intended to lure people somewhere else, and are posted by people who are paid to do that or by bots. It is false to assume that these type of posts and associated accounts cannot be detected or do not exist. I agree that spam by bots differs from spam by humans. However, they need to be dealt with in union because they usually go hand in hand, i.e., bot nets often support and amplify select human posts and the human origins of these networks can be mapped and traced back to particular actors and sources. I doubt anybody would disagree with your general claim that there can be methodological problems with detecting any kind of spam, whether by bots or by humans. Of course, that's difficult, especially for researchers with only limited or no access to sensitive account information such as creation date, post history, IP numbers. If it was easy, then the spam problem would have been solved already, but I have argued that Google hasn't solved it.
- mike_hearn 4y ago"I was referring to posts that are deliberately intended to misinform" But how do you know they're misinformed/intending to misinform, unless you disagree with the post?
- jonathanstrange 4y agoFirst of all, I took you to use "disagree" in the sense of politically disagreeing. You can identify alleged information as being false based on other, more reliable evidence independently of whether you agree or disagree with any political position it is used to support. Second, often the original source can be identified, and in certain cases it's clear when an original source is compromised or unreliable. For instance, sometimes the original source are mere bloggers or activists with no privileged access to information. Other times the source is a known misinformation outlet. For example, wareonfakes.com is registered in Moscow, has no identifiable sources or posters or anyone else taking responsibility, and is obviously not upholding any journalistic standards. Third, it is often forgotten that misinformation can also involve correct information - this is even the standard. The misinformation part is in skewing someone's perception of reality. This can be recognized with a bit of common sense. For example, reports of small, anecdotal incidents are irrelevant for judging anything. The "the lady who defied a soldier from country X" were spread both by Ukraine and Russia, and are irrelevant to evaluating the overall situation. Other reports are not credible from the start. A typical example was the "Ghost of Kiev" story, which doesn't even need fact-checking to tell it's main purpose was propaganda and the informational value was almost zero. Fourth, not all of this is hard to recognize, sometimes misinformation is spam in the traditional sense. For example, there used to be (or still are) a massive number of posts under Youtube comments to advocate the not-yet-closed Youtube channel of a vlogger of US origin who claims to be a journalist (for which the person has no credentials) and is embedded into the Russian military campaign. These kind of repetitive referrals to other outlets are spam under any definition of "spam", and violate Youtube's ToS. Fifth, you can recognize intentions by analyzing the texts and behavior of the actors, and analyzing their amplification networks (who repeats which Tweet, for example). Humans are fairly good at recognizing bad intentions, although it's arguably harder when you only have written text available. A social network analysis helps, since bad faith actors elicit different publication patterns than ordinary users. Last but not least, in the case of Russia's current aggressive war against Ukraine, it's not just about misinformation. Not everyone is aware of that, but in many jurisdictions supporting Putin's war can constitute a crime punishable under penal law, just like using the word "war" can constitute a crime in Russia. In my opinion, all posts in favor of Russia's aggression should at least be removed from social media. This shouldn't even be controversial, 141 countries have condemned this aggression which violates all international laws. It was not common during WW2 for US companies to give the German Nazi party free airtime on all broadcast channels. People could have their personal, pro-Nazi opinions in allied countries, but expressing them publicly had - and should have had - negative consequences. There is no reason to think this should be otherwise in the current conflict, which is essentially a large proxy war between NATO countries + Ukraine as defenders against the sole aggressor Russia. Not just misinformation should be stopped in this particular case, but also open expressions of support for Russia, displaying the swastika-like "Z" sign, and so forth. And, of course, the NSA and other intelligence agencies should hack back to destroy Russian propaganda outlets and take them off the net. However, I admit that this last point kind of deviates from the original discussion. In a nutshell, there are many indicators for spam and many indicators for misinformation, and you judge these in a way similar to how a judge would evaluate cases that primarily involve circumstantial evidence. Humans can do this reasonably well, algorithms are still bad at it.