4 ms·
Spammers are possibly trying to take advantage of npmjs.com domain's high Google rank. I found and reported this spam account [1] with links to download movies.
by sirius87 4y ago
Spammers are possibly trying to take advantage of npmjs.com domain's high Google rank. I found and reported this spam account [1] with links to download movies. They seem to be using npmjs as a free web host with good SEO.
[1] https://www.npmjs.com/~aarilzd https://www.npmjs.com/~aarilzd
- Ciantic 4y agoIf the spammers only want to be indexed, then NPM should disable indexing for major search engines. But still allow it to be indexed other ways, which aren't unearthed on Google search. Other ideas include: do not index new packages before they've garnered enough downloads.
- leshenka 4y agohow do you garner enough downloads without being discoverable by Google?
- Ciantic 4y agoIt's a fair question, most JS libraries I've discovered weren't directly accessed with Google -> npmjs.com but instead from the library's own page, GitHub, Hacker News, etc. If I Google a library and end up on npmjs.com I usually just click on a link to the library's repository or home page first. Of course, it would disenfranchise a bit, but what is another option?
- counttheforks 4y agoNot npmjs.org's problem. Most languages their dependency managers don't give away indexed flashy web pages for free either, yet discoverability is usually not a problem.
- chatmasta 4y agoWhich languages have dependency managers with a public registry that is not indexed in Google? pypi.org and docs.rs are both indexed in Google, for example. With docs.rs it's even kind of annoying because often the indexed page is for an outdated version of the package. There's really no reason why the same spammer couldn't target those sites too.
- prepend 4y agoAs a developer, I want npm package information and docs to show up in search. I frequently prefer pypi or cran results over others because then I can easily tell if it’s a usable package vs just some snippet. Especially cran because it has pretty rigorous entry requirements so being in cran is a signal of at least some minimal quality.
- onion2k 4y agoAs a developer, I want npm package information and docs to show up in search. What case is there when you want to find a package in NPM, and information about that package, using Google? If you want information about the package then it's find if the NPM package page is missing from the results - so long as you're getting the package's homepage or git repo then that's plenty. From there you can get to it's NPM page. If you know the package you're looking for, or if you know what you want to do, then searching NPM itself alone is fine. Essentially, there is no overlap in the Venn diagram of "searching for a package" and "searching for information about a package". You want one or the other, not a results page with links to both. If people realized this about their searches more then Google could fix a lot of spam problems.
- unbalancedevh 4y ago> there is no overlap in the Venn diagram of "searching for a package" and "searching for information about a package". I don't know, if I want information about something, it seems pretty reasonable that I might do my search for that something.
- onion2k 4y agoIf that's the case then you're doing the second of the two searches, and if the NPM package wasn't in the results but its Github repo or homepage was you're still getting the results you wanted. For any search where you don't know what you want Google without NPM pages works fine. For any search where you do know what you want NPM's search function works fine. There isn't a case where you need Google to interleave pages from the wider internet with pages from NPM. You only think you want that because it's what you're used to, or because you use Google to do searches that you should really do on NPM instead.
- loic-sharma 4y agoGoogle search is an extremely common way to discover packages. Disabling indexing entirely isn’t a valid solution. Downloads are very easy to fake. Usually package managers don’t allow indexing until the package and its author reach a certain age. This allows the team to discover and remove the package before it is indexed.
- SquareWheel 4y agoThat seems pretty extreme. Why not just add nofollow to links? That's what websites like Wikipedia do.
- franky47 4y ago> Other ideas include: do not index new packages before they've garnered enough downloads. Which would be trivial to automate.
- marginalia_nu 4y agoAs an aside, something I've seen when reverse-engineering black hat SEO is online casinos sponsoring prominent open source projects in exchange for a sponsorship link. Seems generous until you you realize this also means a huge boost in page rank.
- sirius87 4y agoI've seen this in the Linux Mint project [1] with donations coming from carpet cleaning and light fixtures cos. Sometimes you'll see law firms and I.T. consultants. It's a pretty great idea. Counts as a win-win in my books, as long as the biz is legit. [1] https://blog.linuxmint.com/?p=4466 https://blog.linuxmint.com/?p=4466
- r9295 4y agoWow, I really wonder how people come up with such attack vectors
- dmux 4y agoThey must do a pretty good job of automating the removal of such packages because I get a 404 from that link.
- sebzim4500 4y agoPresumably this will only start to happen more when LLMs are being trained on this kind of data. For example, every training corpus weights Wikipedia way higher than random websites/forum posts, so sticking an ad for your product on some random article that no one looks at will get it into the model.