9 ms·
The New York Times can graph and identify bots, and other publications / bloggers have been able to identify networks of them. Supposedly, the best and brightes
by randomdrake 9y ago
The New York Times can graph and identify bots, and other publications / bloggers have been able to identify networks of them. Supposedly, the best and brightest get recruited by companies like Twitter and Facebook, but for some reason they're incapable of identifying and shutting these things down?
What is all the hubbub about machine learning, and things like neural networks, if they aren't being actively employed by the tech giants?
There's only a couple possible scenarios I can come up with for why this continues to occur:
1) The best and the brightest actually don't work for any of these companies; they're just constantly trying to catch up to teams of developers more highly skilled than them. They are bright, and decent, but ultimately mostly average.
2) The developers in these companies are on par with those who can graph and identify these networks of malicious and fake content makers, but they just don't care because # == $.
3) The companies actually have their head so far in the sand that they don't have the technology, or the resources, to combat this.
I find it hard to believe that people outside of the ecosystem of these companies have more skill, knowledge, or capabilities (1). It also pains me to think that hundreds of millions or billions of dollars simply cannot run a company that can combat this (3).
So that seems to leave complacency (2), which is disturbing and definitely should send up more red flags about whether we should be giving these networks any attention at all.
I'd love to hear other possibilities or maybe more information to back up any of the other scenarios.
- jayess 9y agoIt seems to me the simplest answer is that they're fully aware and not doing anything about it. It doesn't take the best and brightest to figure it out. I've purchased thousands of twitter and facebook followers before for $5 on fiverr just for fun. There's no reason twitter or facebook can't do the same and identify the fake accounts. The logical conclusion is that they don't want to.
- pdx 9y agoI agree. Just buy them. If they spend $100K a year on buying fake accounts and then banning those accounts, that will be more effective than messing about paying a team of 10 engineers to do pattern analysis. If they spent $2M a year on buying and then banning fake accounts, they would squash it completely. They're not doing it because they like the big follower counts themselves. They're in on the con.
- stouset 9y agoHow does paying the botters market rate raise their cost of entry?
- kibwen 9y agoThe price of a single bot is amortized across thousands upon thousands of follows. If it costs a penny to have one bot issue one follow, and a single bot performs 5,000 follows, then that bot has earned $50. If Twitter were to buy follows for the market rate (one penny) then ban those accounts, it would drive the market rate for bots up to (using our hypothetical example) $50 per follow. (This is an oversimplified example, of course, but reducing the average number of follows that a given bot can perform would indeed have the effect of raising the market rate, though we can continue to argue over the magnitude; it's self-evident that if generating bots didn't have fixed cost overheads, then we'd see fewer bots with incredulously inhuman numbers of followed accounts.)
- jaggederest 9y agoThis seems liable to cause unintended side effects... https://en.wikipedia.org/wiki/Cobra_effect https://en.wikipedia.org/wiki/Cobra_effect
- pdx 9y agoI disagree. The cobra effect occurred because the English provided a market for cobras where there was no market before. There's already a large market for followers and twitter adding itself to that market in such a limited way does not add significant new demand.
- AlexCoventry 9y agoThe Reagan administration tried this in the Iran-Contra affair. They only succeeded in incentivizing more terrorism and abduction of US citizens.
- tabeth 9y agoYour premise is centered around tech companies wanting to shut these things down. Is that actually true?
- sixo 9y agoAnd that they want it to a degree that exceeds the cost of doing it, and of keeping up in that arms race. Chances are they don't.
- alextheparrot 9y agoThink of it like cancer. Identifying cancer, saying “You have a growth”, is one problem is medicine. There are ways of doing this and the cancer is just fine being identified, it is none the wiser. The moment you start to treat cancer is the moment you realize that identifying and treating the problem are leagues apart. Cancer mutates, cancer squirms, cancer goes incognito until the storm has blown over. By attempting to treat cancer, we’ve made it our adversary and it now will fight against our treatment. UV treatments cause mutations that the cancer might use to avoid further treatment, some drugs target specific causes of cancer which the cancer will then change to, and some treatments just kill the patient (Just as Twitter can’t just ban everyone). There is a balance that these large tech companies need to find just as a physician needs to determine the correct path. I know this is a very optimistic view of the problem, but it aligns with my experience working, generally, in this space at similar scale.
- revelation 9y agoFor the longest time, spammers and fake accounts on Twitter were just a name with a long ass random number on the end. I guess that particular cancer hit a local optimum where Twitter was pumping its numbers aka misleading the shareholders.
- losteric 9y ago(2) except the blame falls on company leadership - developers are not in charge of funding teams. Leadership funds teams that generate positive ROI. A successful bot-hunting team will likely decrease revenue... although kicking out bots makes the platform more sustainable. It's not just social media. The same question kicks around retail sites like Amazon - why are external researchers better at identifying fraudulent reviews? Because review factories propping up shoddy Chinese imports are very profitable for Amazon. These days customer obsession is just a "nice to have".
- Strilanc 9y agoTwitter has a much higher cost for false positives. If 1% of the accounts this article thought were bots were actually legitimate, no big deal. If Twitter banned that 1%, that's ten thousand pissed off people. Also, bot makers don't care about fooling the New York Times, they care about fooling Twitter. If Twitter applied this analysis, the bot makers would adapt to it. It only looks like it works until you start to use it.
- kibwen 9y ago> If Twitter applied this analysis, the bot makers would adapt to it. The problem with the arms-race argument here is that it implies that just because an arms race is inevitable that the only acceptable course of action is total capitulation. Twitter is bigger than any of the botters, and if it wanted it could raise the cost of entry of botting high enough that nobody would be able to afford the bots in any quantity large enough to make a difference. They just simply don't want to. (As for false positives, I've seen plenty of people whose accounts were suspended for reasons unclear to them; Twitter seems to care very little whether its captive userbase is pissed off or not.)
- kbenson 9y agoI don't think it implies that unless you've already accepted they aren't doing anything. I think it's entirely possible that a lower class of bots are routinely blocked, but a higher class are always mutating to get past the latest filtering, so has a large fairly stable population.
- dmoy 9y ago> Twitter is bigger than any of the botters, and if it wanted it could raise the cost of entry of botting high enough that nobody would be able to afford the bots in any quantity large enough to make a difference. How would they do this without triggering enough false positives to destroy their product? I mean seriously, if you are capable of doing this that well, you could probably walk into a job paying >>$750k/yr at any of the large tech companies, because they'd be falling all over themselves trying to get at your magic.
- tlb 9y ago4) It's not hard to identify bots with 95% accuracy, which is good enough for reporting, but you need 99.9% accuracy in order to delete their accounts or else too many genuine accounts will be wrongly terminated.
- tmalsburg2 9y agoSuper easy to solve: identify likely bots, suspend account, and let owner do something trivial that only a human can do. False positives solved but the cost of maintaining large amounts of bots are prohibitive. Also, high-profile non-bot acounts should be trivial to identify with near certainty.
- jimnotgym 9y agoThey could join the rest of the internet and serve a Captcha to the suspicious accounts!
- greenleafjacob 9y agoThey already do this but instead of a CAPTCHA they require a SMS. I don’t know if they had the option to enter a CAPTCHA instead?
- robryan 9y agoIt shouldn’t be too hard to get at networks. If you identify an account with suspect followers and then look at the followers common follows and follow dates. They can spread out the follows but still to get it done it any reasonable time the follow times are going to be pretty close.
- dredmorbius 9y agoThere are actions other than deletion which can be undertaken. You might target suspect accounts for increased validation -- Captcha or other elements, or simply logging them out more frequently. Changing programmatic elements such that they cannot be scripted as easily (this is one of the primary credible arguments I've seen against APIs for large public services). (Craigslist, and several other services, have applied what's effectively a service-degredation level to accounts suspected of malicious activity. It's not a hard fail, and it's possible for a real human to get past the hassle, but it slows down attackers considerably.)
- mbesto 9y ago> There's only a couple possible scenarios I can come up with for why this continues to occur: Or....this has nothing to do with talent and more to do with market dynamics - they lack any monetary incentive to change.
- colah3 9y agoIt seems to me that there may be a more benign explanation: Facebook/Twitter/etc are in a large-scale, iterative, adversarial game with many opponents. It might be relatively easy for humans to catch clusters of fake followers, but that doesn’t scale. If you try to create heuristics to catch the fake folllowers, the adeversaries will try to adapt. If you try to learn rules, you can update your heuristics faster, but you need trading data, your adversaries could learn too, and adversaries could try to do data poisoning attacks. It seems like a really tricky situation for Facebook and Twitter.
- flatline 9y agoI think that you are right, but only in the qualitative sense, not the quantitative. Spam networks have shown us that taking out a few big players - the largest few sub-graphs of the network - can have dramatic effects on spam rates, reducing them by e.g. a factor of 20. I’m confident that Twitter could identify and eliminate a large percentage of the problem overnight. So far they not even acknowledged that there is a problem. O.P. asked for alternative explanations and people are trying to give these companies the benefit of the doubt. There are actually some much less generous explanations, such as that they are knowingly in the pay of forces adversarial to US interests.
- 18pfsmt 9y agoMy belief is that they are still profiling their spammers, and are not acknowledging anything specific as a sort of 'poker face.' >people are trying to give these companies the benefit of the doubt. I think that's a good thing; we're all too susceptible to cynicism. Being charitable is perfectly fair (and ideal) when dealing with people who haven't demonstrated a propensity to deceive (or operating with malicious intent).
- yuliyp 9y agoHow do you "take out" an attacker? They will just create more accounts. The accounts themselves don't have to cluster with each other. And let's say you do find a large set of accounts and ban them. New fake accounts are being constantly created and sold and used.
- flylib 9y agoI mean prolly somewhere near 40/50% of Twitter's 330 Million MAU are fake, the company will drop 50% or more in value if they ever decide to delete those accounts, incentives are aligned for Twitter to encourage bots by making it easy as possible (not adding captchas, etc.) and not enforce policy
- DonHopkins 9y agoIt would be awesome if the New York Times shared the tools they've developed for this article to identify and analyze bots! I usually hate those scroll-responsive animated web pages, but the scrolling illustrations and data visualizations in this article were particularly well done and pretty amazing. Not just pretty (the initial face sequence) and clever (the scrolling iPhone) but also actually useful and relevant (like the subscriber graph data visualizations expanding over time). I would love to know more about the tools they used to make those too. A plea to NYT from a subscriber: Please share some of that great software you've developed, and publish it on your github account!!! https://github.com/NYTimes https://github.com/NYTimes
- rich_harris 9y agoInvestigations graphics editor here — as it happens a lot of the tools we use are open source! The interactive components were all built with Svelte (https://svelte.technology https://svelte.technology) and Rollup (https://rollupjs.org/guide/en https://rollupjs.org/guide/en). We'd like to eventually open source some of the stuff we built to do the analysis as well, though it depends on time and priorities.
- vthallam 9y agoHey Rich! The graphics are on point for this article. I really liked the unfading faces at the start. Also, the times should put a tip jar on these kinds of articles, I'm not an avid reader to subscribe, but would want to appreciate any good articles I find.
- DonHopkins 9y agoThank you for those links, and your great work. If Twitter refuses to do the kind of investigation and analysis that you performed because they're making too much money from bots, then the free press and open source community needs to take up the slack!
- catacombs 9y ago>it depends on time and priorities. AKA never.
- tw1010 9y agoThe work of finding exploits to these algorithms probably requires less talent than the work of researching and implementing them.
- avip 9y agoAll relevant parties in FB, twitter, Yelp, TA are very much aware of fake accounts, fake review, and any other bot-related traffic. Yes, they do have teams dedicated only for that. So the real question should be - what actions are being taken to counteract social bots. This is not something any company would want to go public about. So basically - we don't know.
- phjesusthatguy3 9y agoQ: What do you get when you cross the Weekly World News with a Baleen whale? A: The American electorate.
- mschuster91 9y ago> The New York Times can graph and identify bots, and other publications / bloggers have been able to identify networks of them. Because it takes a lot of manpower to do so on a wide scale and almost always raises the question of fairness - the NYT (and others) focus on the egregious examples, but if Twitter would do the same (and kill the followers) their customers (the big celebs) would complain, and when they say "okay, bye Twitter", the normal users who want to follow them would disappear... thus reducing the popularity of Twitter for advertisers. The bots don't count as advertisement recipients, so Twitter doesn't have advertisers complaining about fraud from these fake followers. In addition, the existing (moderation) manpower is focused on keeping the system being overrun by actual, human-powered Nazi and other troll accounts - fake followers are waaaay on the bottom of the priority list. They don't generate bad press headlines about shitstorms or harassment.
- sgk284 9y agoI used to work at Twitter, so not sure I can get into specific numbers, but they remove hundreds of thousands to millions of accounts every day (literally every day). It's particularly tricky because bots aren't inherently against the ToS (e.g. the earthquake bot cited in the article), so you can't just ban based on "Is this account behaving like a human?". Many "bots" though are actually giant networks of humans being orchestrated and they share things that are compelling to a subset of real users (e.g. Hillary did this bad thing... lock her up), so now real users are retweeting and muddying any signal that was there to detect a bot network. It's an incredibly challenging problem at scale and you basically only see the < 1% that Twitter missed. They've got a pretty sizable team that works on it so hopefully in a few years they'll have solved it much like Gmail has mostly solved spam.
- djsumdog 9y agoI think it was the 99% Invisible podcast that did an interesting report about whole networks of people who were paid in Mexico to spread election messages; and these networks are often reused by gangs or other organizations to promote stuff or harass opponents. So in this case, they're not bots. They're actual humans, because it's cheaper to do that than pay people to write scripts.
- amsilprotag 9y agoI don't remember that 99pi. Was it Reply All? https://gimletmedia.com/episode/112-the-prophet/ https://gimletmedia.com/episode/112-the-prophet/ After Andrea is attacked by a stranger in Mexico City, she just wants to figure out who the guy was. Investigating this question drops her right into the middle of one of Mexico’s biggest conspiracies.
- djsumdog 9y agoAh yes, sorry you're right. It was Reply All.
- kakarot 9y agoAdditionally, you can train a human to become better at astroturfing over time, and train your network as a whole to be resilient to partial takedowns. These things are much easier to orchestrate than bots, especially if your specialization is not CS but scamming / manipulating opinions, and in many cases, especially with scale, human labor can actually be more expensive. Of course, subsidizing the labor to impoverished countries definitely helps; these networks tend to lack the sophistication used in more sensitive political or industrial topics. Sometimes for national or global campaigns you'll find a mix of both cheap labor, a small team of experienced astroturfers with a good grip on their persona acquisition and management, and bot networks, each targeting a different set of demographics. If your target demo is sufficiently ignorant or radicalized in their beliefs, it doesn't take much to convince them and bots will do just fine.
- modeless 9y agoYour argument that NYT is better at finding bots only makes sense if you assume that NYT found all the bots that exist. A more likely explanation is that everyone who goes looking for bots uses a different methodology and finds a different subset of the bots that exist. Thus NYT can find some bots that Twitter missed, especially by using labor intensive methods that don't scale to all of Twitter, but there are plenty that NYT missed too. Their methods aren't necessarily better than Twitter's, just different (and likely much more expensive). Note also that the NYT people don't have to care about their false positive rate and don't have to adapt their methods to adversaries since this is a one time analysis.
- pryelluw 9y agoThis is not a technological challenge, but a management one. Why would a product built to expand social graphs delete nodes that further it's size? Not gonna happen. Ever.
- anametoremember 9y agoThe object of a political botting campaign is to promote views that are complementary to ones' own political agenda and remove those that are not. With an automated system that removes discourse that corresponds to patterns used by bots, a social adversary would only have to adjust their bots to resemble the voices they want to silence, for a perfectly platform-sanctioned way to silence people they disagree with.
- alexbeloi 9y agoFor (2), developers aren't the ones making the decision to deprioritize bot detection. The decision makers don't care because the ROI on detecting/removing bots is likely negative, whereas using those ML people to improve ad targeting and engagement is much higher.
- unreal37 9y agoYeah exactly. All the devs are working on business prioritized projects. Removing bots is not a business priority. Nothing to do with how smart the devs are.
- jayd16 9y agoIt's not in the company's interest to ban these bots but if it really is that easy to track, we should do it from a third party. A Twitter/Facebook/whatever client plugin that flags or hides bots seems feasible. It's even in a certain business interest to track bot traffic for more accurate click through tracking so perhaps there's even a market need for this.
- AlexCoventry 9y ago> I find it hard to believe that people outside of the ecosystem of these companies have more skill, knowledge, or capabilities The methods described in this article would be corporate espionage. I don't know what the legal implications would be, but ethically, I'm personally a bit frightened by the idea of a company as large and in possession of so much personal information as Twitter getting into the espionage business.
- vadimberman 9y ago4) The tech is yet to achieve acceptable level of accuracy.