7 ms·
Is it very difficult? So many of these spam accounts follow the exact same format, with names following some variation of "Free stuff Telegram: +1234…" or "s3xY
by thamer 4y ago
Is it very difficult? So many of these spam accounts follow the exact same format, with names following some variation of "Free stuff Telegram: +1234…" or "s3xY p1cs c0mE 2 mY chAnNeL", etc. Tons of them post these obviously made-up threads about how "Dr John Smith" has helped them triple their crypto investments.
There's probably a dozen or so templates that many of these spammers and scammers use, and if you can spot them in a second, it's not hard to train a model to recognize them also.
It is out of control because YouTube does not care to fix it, not because it's some insurmountably hard problem that one of the top AI companies in the world just can't figure out.
- chx 4y ago> if you can spot them in a second, it's not hard to train a model to recognize them also. something Minsky vision something (I know it's an urban legend but still, it figures.)
- meheleventyone 4y agoI can totally see that if you assume it’s simple and easy to fix that it must just be another case of people not caring. Yet here we are replying to a post about the not-ideal steps the company is taking. And further this is a problem for multiple large companies with user commenting and chat, from Google through Roblox, Twitch, Twitter and beyond. There are a few things that make naive solutions impractical: * Cost to build them. * Cost to keep them up to date in a highly adversarial environment. Note this might mean scraping the solution and starting again. * Cost of running them at huge scale. Particularly important is to understand the attack surface is so high and cost to the opposition is so low that people will defeat your approach recreationally just to spam obscenities.
- zasdffaa 4y agoGuy makes concrete suggestions, you respond with abstract rebuttals that add up to "it's too hard". All he's saying is low-hanging fruit is pluckable; pluck it.
- meheleventyone 4y agoGuy doesn’t understand why perceived low hanging fruit hasn’t been plucked and I explained why it’s not that simple or low hanging. That’s fine, naive filtering (now with added ML!) is the first idea everyone comes up with. Then they learn about the Scunthorpe problem and beyond.
- deleted 4y ago[deleted]
- doublerabbit 4y ago> Guy doesn’t understand why perceived low hanging fruit hasn’t been plucked and I explained why it’s not that simple or low hanging. I'll explain too; Money. YouTube make money from the spam accounts, As twitter does, Reddit too. Spam is a very profitable market.
- smallnix 4y agoHow does YouTube make money from spam comments?
- doublerabbit 4y agoPay Google, Google turns a blind eye. Why else has Google not cracked down on spam accounts? It has the tech, it knows it has a problem and the best remedy they can come up with is shortened YouTube usernames? There's more to it behind the scenes. A real example is that I worked for a business where my job was to sell email addresses gathered from soft-core pornography websites. A batch of working emails could easily fetch £500. Reddit will sell you accounts for your "campaign" if it profits reddit and it's all the same for any big walled garden.
- Nextgrid 4y agoSpammers contribute to user & engagement numbers and might even drive real engagement behind themselves as (angry) real users respond to them.
- franga2000 4y agoThere is a community project out there to filter and delete most of the spam and many creators use it. It was developed primarily by one person in their spare time. Google has no excuse.
- saurik 4y agoTo be fair, Google has tens of thousands of developers, and so would have a hard time being anywhere near as productive as one person.
- joshspankit 4y ago^ this might sound like a joke to many, but is hauntingly accurate.
- MichaelZuo 4y agoThe common view on HN seems to be that google products optimize for promotions, and not even executive or entry level promotions, primarily middle management promotions. I don’t think it’s entirely true as hopefully the finance department has folks smart enough to realize if headcount is adding negative value or positive value.
- numpad0 4y agoI wonder if the problem is sheer technological difficulty or if it's similar to antibiotic controls to prevent resistance development.
- paulpauper 4y agoThey don't even need to change the name. I recall seeing a giveaway livestream scam in which the hacker didn't change the name of the channel at all and it was still bringing in btc.
- thrown_22 4y agothаmеr is not thamer. The space of all unicode characters is too big for you to look for every single possible "looks like" pair. Unicode, as unpopular an opinion as this is, should only be used for presentation and not storage. Anything past ascii is a security vulnerability.
- jfim 4y ago> Unicode, as unpopular an opinion as this is, should only be used for presentation and not storage. Anything past ascii is a security vulnerability. While at it, we should also ban the letters I and L as they can be confused. People who have those letters in their name can just request a name change. /s Leave Unicode alone for the vast majority of the world whose language isn't actually writable using ASCII.
- thrown_22 4y agohttps://en.wikipedia.org/wiki/File:QWERTY_1878.png https://en.wikipedia.org/wiki/File:QWERTY_1878.png Notice some keys missing there? The difference is that we have a few good quality fonts, like liberation mono, which let you tell the difference between l and 1 and 0 and O. There can't be a font which does that for the whole of unicode space because a and а are by definition the same glyph.
- jfim 4y agoWhat's your proposal for differentiating between mais (but) and maïs (corn), or between café (French, Spanish) and cafè (Catalan) using ASCII? How would one write an ideographic language using ASCII?
- iforgotpassword 4y ago> What's your proposal for differentiating between mais (but) and maïs (corn), or between café (French, Spanish) and cafè (Catalan) using ASCII? Context, but... > How would one write an ideographic language using ASCII? ... We're talking about security sensitive context. Users can write stories with all the Unicode they want, just don't use it for identifiers like user and channel names. Reddit got it right.
- ehnto 4y agoIt is an adversarial space, the reason the names look ridiculous is because they have evolved to escape pattern matching as it also evolved. YouTube and Google are notoriously unhelpful when it comes to instances when they catch a human in their automation drag nets. They lack the capacity to support their platforms so they are probably being cautious not to create support and ill will they can't handle.
- nikanj 4y agoIt’s hard, if you’re insisting on doing it the Google way i.e. ML+AI with zero humans involved. Having a person keep an eye on the top 50 popular channels and spotting the patterns in commentor names would be very effective, but having a human is not Google
- iforgotpassword 4y agoExactly this is my guess too. Literally everyone is seeing these very obvious spam comments everywhere. Have a few people monitor comment sections manually and then maybe one guy who goes over the spam ones and come up with a regex-like matching rule to detect these and move them to the "potential spam" section of the creator's dashboard. It's gonna be whack-a-mole but it's rather responsive. I get it, they want to automate everything, but evidently that only gets you so far, your model will never be at 100%, and we didn't even get to false positives yet.
- MauranKilom 4y agoThat means you're now selecting for all the spammers that are somehow evading the eyes of that crew. It doesn't matter whether they are doing it intentionally initially, this un-natural selection will inevitably lead to spammers figuring out how to not be seen by those Google people.
- jfoster 4y agoIt sounds like you're assuming that Google sticks to a single approach and doesn't waver from it, despite spam still being present. They wouldn't do it like that, though, would they?
- MauranKilom 4y agoThe point is more that, unless Google keeps continuously expanding their approach along every dimension, whichever spammers happen to be evading them at any given point will stick around and multiply.
- bryanrasmussen 4y ago>There's probably a dozen or so templates that many of these spammers and scammers use, and if you can spot them in a second, it's not hard to train a model to recognize them also. I'm betting it's probably not hard the spammers and scammers to switch templates if old ones get caught.
- norman784 4y agoI think a company like Google have the resources to dedicate a team to work on that permanently, but as I suspect, working in something like this doesn't get you enough merits, so it's better to work on a new shinny useless feature than an actual useful but not shinny feature.
- can16358p 4y agoSame for Instagram and Twitter. The formats of the obvious spam is super simple to detect on both platforms for years. Yet I haven't seen ANY progress on either platform. Assuming that they are much more competent than a random dude like me as a whole tech company, the only possible explanation is that they have incentive not to prevent spam. Perhaps: spam = more notifications delivered = more app opens = more potential feed engagement since they've opened the app even if from a spam notification = more ads shown = more revenue = happier shareholders short time... all at the expense of losing long-term trust to the platforms.
- ElemenoPicuares 4y agoHave you been working on actually detecting them or are you assuming they’re easy to detect? Did you go digging for false negatives or false positives in other languages? Considering that not all special characters are printable, these platforms don’t limit their users to ASCII, random foreign words are common in other languages, the false positive rate has to be pretty much zero, etc etc etc I’ll bet it’s a whole lot more sophisticated than you imagine.
- can16358p 4y agoI'm assuming they're easy to detect yet I'm super confident about the patterns I'm seeing that I'd bet money that it can be implemented. False positives should ideally be zero yet algorithms don't have to ban right away. If there's a relatively small team of people reviewing those spams manuallu after the algorithms flag, I don't see any problems with the approach. I of course get there are many different corner cases, Unicode characters that look like other letters etc. to bypass the filters yet they're also very easy to implement a filter that would probably detect many spams right away even with many of the corner cases. But again, a human team would still review these tweets anyway giving no room for false positives.
- ratww 4y agoAbout 90% of the public Youtube spam I see on YT still follows this ultra-simple template: one person asks something on the "top comment", another one replies with a faux-legitimate message. The question is often "does anyone know how to pirate videos?", "...how to invest in crypto", "...how to get fake instagram followers", life coaches, etc. It's not really custom messages with special chars, it's something that I could detect by using Ctrl+F in the browser. It's not something they'd need special tools on their end to catch. The only thing that is changing is the name of the users. So yes, I'm gonna agree with GP that "the formats of the obvious spam" are super easy to detect.
- deleted 4y ago[deleted]