4 ms·
I don't see anything positive coming out of academia having access to the Twitter firehose.
by favaq 4y ago
I don't see anything positive coming out of academia having access to the Twitter firehose.
- makestuff 4y agoWhen I was in college we used it to try and try a sentiment analysis model since they are notoriously bad at detecting sarcasm and Twitter was full of sarcasm. We also used the API to try and determine the most impacted areas after a natural disaster. Basically it would use the model we trained to try and read tweets of people that needed help or people tweeting about severe damage and group them by their coordinates. The first one I agree isn’t really positive since it is just using other people’s data to train a model, but the second one could’ve been a useful tool to help EMS during a natural disaster.
- isubasinghe 4y agosigh this is just straight up wrong, I was an RA that worked on a real time social media analytics software. We were able to pick up on things like likely covid infections sites etc.
- mike_hearn 4y agoTry searching Google Scholar for "social bot", or to save time, just read this paper: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3814191 https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3814191 Academia has flooded the literature with >10,000 research papers based on the Twitter API feed. Virtually none of it is reproducible, it's frequently based on circular logic, the methodologies are unscientific and the conclusions are usually deeply partisan, but it nonetheless gets amplified by the media as "proof" of various false claims. Count me in the camp of people who is happy Musk is doing this. I've been writing for years about the plague of "social bot" research coming out of academia that's based on the Twitter API: https://blog.plan99.net/fake-science-part-ii-bots-that-are-not-c66129e5e3f5 https://blog.plan99.net/fake-science-part-ii-bots-that-are-n... https://blog.plan99.net/did-russian-bots-impact-brexit-ad66f08c014a https://blog.plan99.net/did-russian-bots-impact-brexit-ad66f... Maybe your specific work on COVID was good, but it was certainly drowned out by the work that was sharply net negative for both society and science. Academic institutions were clearly never going to get the problem under control, so booting them out whilst allowing search engines and the like to continue accessing the feed seems like a good solution.
- anigbrowl 4y agoThis is absurd; you're throwing the baby out with the bathwater. Certainly, it is easy to find social science papers with terrible methodology that use the Twitter API, or that build on the sand of papers with terrible methodology. But you conclude from that that all academic use of the Twitter API is garbage, which is nonsensical, and that preventing academics from studying Twitter at scale is the ideal solution. Your hyperbolic language (here and in your two medium articles, which I read thoroughly, along with the SSRN paper you cited*) does nothing for your own credibility. The main 'methodology' of the SSRN paper is combing through other papers' datasets, contacting some of the identified 'bot' accounts, and establishing that they're operated by real people; the accounts as misidentified as bots when in reality the account operators were just aggressively quote-tweeting by using copy & paste to spread (eg) political or Qanon messages 200 times an hour. The authors point out that by really making an effort, Twitter users can tweet spam up to 25 times a minute, with no bots in sight! While the authors are quite correct to point out that people can be misidentified as bots, this completely ignores the fact of the unwanted spamming behavior. Pointing out the scientific flaws of 'tools' like Botometer is wholly valid, but the effort to research and develop tools for bot identification are a response to the fact of systematic information pollution, and most papers that try to address this issue are careful to offer caveats and qualifications about the limitations of their methods. It is not the fault of academics if media pundits over-simplify the fruits of their research. Here are some examples of high quality research using data from Twitter: https://www.researchgate.net/publication/336638958_Ephemeral_Astroturfing_Attacks_The_Case_of_Fake_Twitter_Trends https://www.researchgate.net/publication/336638958_Ephemeral... https://www.researchgate.net/publication/334816353_Political_Astroturfing_on_Twitter_How_to_Coordinate_a_Disinformation_Campaign https://www.researchgate.net/publication/334816353_Political... https://www.researchgate.net/publication/361949311_QAnon_Propaganda_on_Twitter_as_Information_Warfare_Influencers_Networks_and_Narratives https://www.researchgate.net/publication/361949311_QAnon_Pro...
- mike_hearn 4y agoI'll try and find time to look at the papers you cite as high quality later today. > you conclude from that that all academic use of the Twitter API is garbage, which is nonsensical "All" no, a vast amount of it, yes. Is it nonsensical? Twitter themselves concluded this exact same thing even before Musk, both in public blog posts and internal emails (see the Twitter Files for examples). But we don't really need to cite Twitter as an authority here. Just try to answer this question: what mechanisms exist that are stopping bad science outside the field of social bot research, and why have those mechanisms failed within it? It can't be peer review, university hiring committees and so on because those are all existing within social studies as well. > Your hyperbolic language ... What language do you think is hyperbolic, exactly, and why? > Pointing out the scientific flaws of 'tools' like Botometer is wholly valid, but the effort to research and develop tools for bot identification are a response to the fact of systematic information pollution This is exactly the sort of problem I'm talking about: this justification is circular. We do bad bot research because we know there are bots, we know there are bots because we do bad bot research. If there were actually big problems with social bots then it would be easy to find them and research them; we wouldn't see this situation where basically all papers are seeing patterns in noise. Botometer is a good example of that. You admit that it's "scientifically flawed" but with respect, that language is not "hyperbolic" enough. It's not merely flawed, it's outright useless. It had an FP rate of 50% when tested against a known human dataset. Yet the Botometer paper has been cited over 900 times now (up from ~700 when I previously wrote about it). When exactly does the rest of the world get to call time on this bad behavior by the academy? These people are changing the opinions of world leaders on the back of misinformation, the exact problem they claim to be fighting. > It is not the fault of academics if media pundits over-simplify the fruits of their research. It wasn't media pundits that made academics cite the Botometer paper over 900 times, or write outright deceptive papers like the one I reviewed. The problem here is academia and the institutions need to start taking responsibility for it. Otherwise you're going to get situations like this one: academia will just get cut off from data. People don't have time to try and figure out which little subsections of the academy are following the rules to separate them from the rest.