23 ms·
Botspam apocalypse
- pphysch 4y agoThe only real solution to the abuse of anonymous protocols is to stop using anonymous protocols and use protocols where clients can be held accountable. But that's politically nonviable in the West.
- JZerf 4y agoI don't see any good reason why people can't be allowed to remain anonymous while still allowing website operators from taking measures to stop bot abuse. CAPTCHAs can already stop many bots. Other commenters have also mentioned that things like Proof of Work systems and micro-transactions could also stop bot abuse. These don't necessarily require giving up anonymity.
- pphysch 4y agoIt's not just bots, it's troll farms as well which are "real" people destroying public discourse in bad faith.
- JZerf 4y agoWebsite operators could still take measures to stop abuse from troll farms as well while still allowing people to remain anonymous. A website operator like Twitter for instance could perhaps require users to make a small micro-transaction before allowing someone to make a post. Some equilibrium for the cost of a post could probably be found where most legitimate users would still be willing to pay that cost but most troll farms would not.
- pphysch 4y agoAre you serious? The problematic troll farms are the ones backed by states and multinational corporations. Gating speech behind money only makes the problem worse. The correct approach is to deanonymize reasonably "public" online behavior. This is the only way to hold abusers accountable, and, ironically, democratize free speech. 1 person, 1 voice. Not 1 rich person, 100 troll accounts.
- JZerf 4y agoYeah, I'm serious. Even with Twitter, for example, currently allowing accounts to be created and posts to be made for essentially free, real accounts and posts still outnumber those of troll farms and bots from what I've seen. If those troll farms and bots actually had to pay, I imagine there would be far less. I also imagine that those troll farms and bots are less influential than real people. I believe that the endgame is that if a website operator takes enough measures to stop troll farms and bots, the operators of those troll farms and bots will eventually run out of resources and be forced to curtail their activity. You're right that gating speech behind money could potentially be bad and make problems worse but I only offered that as one suggestion. Instead of or in addition to using money, you could perhaps make a system that uses some type of karma/reputation for instance. Those could still be done anonymously.
- 16amxn16 4y agoThanks god it's not. The right to anonymity is something we shouldn't lose.
- paulmd 4y ago> They're a major part in killing off web forums, and a significant wet blanket on any sort of fun internet creativity or experimentation. > The only ones that can survive the robot apocalypse is large web services. Your reddits, and facebooks, and twitters, and SaaS-comment fields, and discords. They have the economies of scale to develop viable countermeasures, to hire teams of people to work on the problem full time and maybe at least keep up with the ever evolving bots. This is not true at all. There are web forums that are not "web-scale" and don't spend all day fighting bot spam. The solution is real simple: it costs 10 bux to register an account, if you're a nuisance your account is banned and you pay 10bux to get back on. Even the sites that don't require payment for explicit registration - often succeed by gating functionality or content behind paywalls. Requiring a "premium membership" to post in the classifieds forum is an extremely extremely common thing on small interest-based web-boards (photrio, pentaxforums, homebrewtalk, etc). That income supports the site and supports the anti-bot efforts as a whole. The customer isn't advertisers - it's the community itself, and you're providing the service of high-quality content and access to people with similar interests. You need to bootstrap a community first, of course, but it doesn't need to be a large community, just a high-value one. The twitters and facebooks of the world just don't like that solution because they value growth above all other considerations. They'd rather be kings of a billion user website with 200 million bots than a 1k-100k user forum with 100% organic membership and content. And they value engagement over content quality, which is the entire reason comment-tree/vote-based systems have been pushed heavily over web-1.0 threaded forum discussions as well. This botpocalypse is the inevitable outcome of the systems that social-media giants have created, not inherent outcomes of the internet as a whole.
- seydor 4y agoAlternatively, allow $0 signups but approve every new account. It's rather easy to spot spam signups
- Nextgrid 4y agoThere are (at least) 2 kinds of spam - "technical" spam such as bots hammering the web service with requests and consuming resources, and the commonly-accepted definition of spam where bots post promotional or other obnoxious content. I feel like the article here talks more about the first kind. I do agree with your solution for the second kind of spam though.
- echelon 4y agoI thought this article was referring to the upcoming deluge of GPT-3/DALL-E bots that will eventually flood all of online discourse. And whatever future models that will be even more indistinguishable from people - perhaps even ones that are good at "signup flow". That's going to be way worse for humanity than spiders and automated scripts sending too much traffic. This article isn't imagining apocalypse creatively enough.
- Avamander 4y agoWe're certainly heading towards a scenario where internet abuse (due to poor regulation against it, IMHO, it's digital pollution) becomes enough of a nuisance to require increasingly intrusive verification. Though we can all work against that by securing our own systems and preventing them from being abused. Used or unused domains should have a strict SPF policy, website registration (or newsletter signup forms) should have captchas, comments should have captchas. Wordpress or other CMS's plugins should be up-to-date and so on and on. Work on requiring 3DS everywhere, everything in-depth. That way malicious actors would be limited to the services they pay for and that makes their life significantly harder.
- api 4y agoWe're coming up on the end of open forums and open social media. Everything will require intrusive verification. Anonymous forums could exist but they'll require something else like an anonymous payment, a ton of proof of work on your local machine, etc. to filter out crap.
- x-complexity 4y ago> The other alternatives all suck to the extent of my knowledge, they're either prohibitively convoluted, or web3 cryptocurrency micro-transaction nonsense that while sure it would work, also monetizes every single interaction in a way that is more dystopian than the actual skull-crushing robot apocalypse. In the interest of practicality: There's a way to go the web3 route without being laden with transactions: - Mint a fixed-cost non-transferrable NFT to an address, with ownership limit of 1 per address. - Use SIWE (sign-in with Ethereum) to verify ownership of address & therefore NFT. - If malicious behaviour is detected, mark the NFT as belonging to a malicious actor at the server's end & block the account. - Require non-malicious-marked NFTs in order to use the site/app. At most, the user only had to perform 1 transaction (minting the non-transferrable NFT) on any blockchain network where the contract resides, & the costs to do so can be made cheaply with Layer 2 networks. (Polygon PoS, Arbtirum, Optimism, zkSync 2.0, etc) Can this be done entirely without web3? Yes, but the added friction imposed onto malicious actors to generate new addresses & mint new non-transferrable NFTs increases the costs for them considerably. > If anyone could go ahead and find a solution to this mess, that would be great, because it's absolutely suffocating the internet, and it's painful to think about all the wonderful little projects that get cancelled or abandoned when faced with the reality of having to deal with such an egregiously hostile digital ecosystem. In all honesty, there's no perfect solution, just hard-to-make tradeoffs: The prevention of botspam inherently requires tracking in some form to resolve said issue, as there's no immediately-recognizable stateless solution for botspam tracking. Someone has to do the tracking to prevent botspam, which inherently involves in state being changed in order to mark an actor as malicious.
- me_again 4y agoThat seems "prohibitively convoluted" to me, if nothing else.
- birracerveza 4y agoBecause you don't have experience with it. There's nothing complicated about SIWE, minting an NFT and checking its validity, certainly not to describe it "prohibitively convoluted" aside from being scared of web3 keywords. Come on now. Not commenting on op's solution's validity or effectiveness, just replying to your comment.
- hamilyon2 4y agoI experienced this firsthand with government immigration websites. The thing is there are only so many time slots and and people are forsed to use a certain web site to apply, so everyone is hunting for available time and generally none are available. So, some creative people set up bots which check periodically for them. They are paid services which will do that for you. Now we have bots hammering gatekeeper's website. Perhaps hundreds of bots. Which results in the website is being unavailable, serving a serious qps to bots. I think it is only a matter of time before someone will write bots that will apply to application bots hoping that more entries with their information will provide them with better probability of success. This is so dystopian and cruel to the average person, and I don't think there is a good solution besides a deep anti-bot expertise whithin the primary website development team
- TomK32 4y agoI faced a website like this recently when booking a slot at my own German embassy, went a different route around the embassy instead. What I don't like about the slot system: You won't get a convenient time slot anyways so why do they bother setting it up like this in the first place? Why not just register with your contact details and receive an email with a guaranteed spot at a selection of three days instead. No more need to reload and no need for bots. The Upper Austrian government did this for covid vaccinations in the early phase and it worked very well that way (early, high vaccination rate amongst the elderly and certain professions).
- luckylion 4y agoThat reminds me of the chaos that ensued in my state in the early days of Covid-19 vaccination when they were still having centralized systems where the elderly could book an appointment. Of course, they had way more demand than supply but still insisted on First Come, First Served, so you ended up with every member of the extended family being asked to try and book a slot, quickly overwhelming their booking systems. At least you didn't need to worry about it all day: there wasn't a chance in hell to get something 3 minutes after the booking system opened each day. Friends described how stressful it was for them, their parents being completely helpless and essentially fearing that their health depended on getting one of those elusive appointments, and being devastated each day they didn't succeed.
- jart 4y agoThis kind of botspam is usually pretty easy to address with redbean using the finger https://redbean.dev/#finger https://redbean.dev/#finger and maxmind https://redbean.dev/#maxmind https://redbean.dev/#maxmind modules. The approach I usually recommend people isn't so much ip reputation, which can be unfair, but rather it allows you to find evidence of clients lying to you. For example, if the User-Agent says it's Windows with a language preference of English, but the TCP SYN packet says it's Linux and MaxMind says it's coming from China, then that means the client is lying (or being MiTM'd) and you can righteously hellban once it's fingered for the crime.
- unglaublich 4y agoWhat keeps bots from just fixing their acts and reporting correct info instead?
- jart 4y agoWhat's stopping them from clubbing you with a monkey wrench? With bots, to answer your question, it'd probably take another standard deviation in the IQ of the person using it. So you've ruled out all the script kiddies in the world by default. The purpose of this game isn't to have a perfect defense, which is impossible, but rather to make the list of people who can mess with you as short as possible.
- krageon 4y agoNothing. Once this sort of fingerprinting becomes common the common bot frameworks will bypass it.
- elias94 4y ago> has been upwards of 15 queries per second from bots What type of queries are they generating? For what purpose are querying Marginalia? Scraping and filling internal search engines? > If anyone could go ahead and find a solution to this mess I would maybe trying to investigate why are querying your search engine. Is for the search results? Maybe from there you can create and sell an API service. Is for the wiki? Is for research purpose? I would love to see some data, raw or with some behavior derived from it.
- marginalia_nu 4y agoMost of the queries don't seem to be tailored toward my search engine, they're ridiculously over-specified and typically don't return any results at all. As I've mentioned in another comment, my best guess is they're betting it's backed by google, and are attempting to poison their search term suggestions. The queries I've been getting are fairly long and highly specific, often within e-pharma or online casino or similarly sketchy areas. Like > cialis 50mg online pharmacy canada price Either that, or nonsense like the below, where they appear to be looking for CMSes to exploit (although I don't understand the appendage at the end) > "Please enter the email address associated with your User account. Your username will be emailed to the email address on file." Finestre Antirumore Torino > affordable local seo services "Din epostadress delas eller publiceras aldrig Obligatoriska flt r markerade med" > "You are not logged in. (Login)" Country "City/Town" "Web page" erst Point is, none of these queries actually return anything at all. I don't offer real full text search, for one. And the queries are much too long.
- TekMol 4y agoCrypto currency mining could be the solution. If one request to the site generates more revenue than it costs in resources, the bot problem is solved. The author says that he is getting 15 bot requests to his site per second. That is about 36 million requests per month. How much does it cost to serve those? $1000 would seem high. $1000/36M = $0.00003 per request. How long would a crypto currency, that is suitable for mining in the browser, need to be mined before $0.00003 is generated? If it turnes out it is a few seconds or so, the solution would be nicely user friendly. A few seconds of CPU time for access to the site. No ads needed to finance the site. It is kind of telling, that Bitcoin started as a spam blocker. The original "hashcash" use case was to use proof of work to prevent email spam.
- GeckoEidechse 4y agoAs much as I hate the whole cryptocurrency hype myself, I think I agree that a proof-of-work requirement on spam detection that pays in the hosts favour could help solve spam to some degree.
- endgame 4y agoBefore bitcoin, there was hashcash, which aimed to do exactly this: http://www.hashcash.org/ http://www.hashcash.org/ . The original bitcoin paper cites it, in fact.
- kragen 4y agoSatoshi Nakamoto almost certainly isn't Adam Back. It might be enough for the request to require more resources from the requestor than from the server, even if it doesn't actually give the server any money. I mean the requestor probably isn't going to be willing to dedicate more hardware "horsepower" to taking your search engine down than you are to keeping it up. That was the idea behind Hashcash. As for coins, the current Bitcoin hashrate is about 200 exahashes per second, down from a high of over 250 a couple of months ago, and the block reward is 6.25 BTC until probably June 02024. At a price of US$24000/BTC that's US$150k per block (plus a much smaller amount in transaction fees) or about US$1.25e-18 per hash. So your suggestion of US$3e-5 would require about 2e13 hashes. https://en.bitcoin.it/wiki/Non-specialized_hardware_comparison#AMD_.28ATI.29 https://en.bitcoin.it/wiki/Non-specialized_hardware_comparis... says an overclocked ATI Radeon HD 6990 can do about 800 megahashes per second (8e8) so you're looking at about 3e4 seconds of compute on that card, about 8 hours. Maybe one of the altcoins that uses a hash function with a smaller ASIC speedup would be a better fit, although I don't know enough about mining to know if there are any where GPUs are still competitive. Still, it seems like it might be more than a few seconds?
- Nextgrid 4y ago> The only ones that can survive the robot apocalypse is large web services. Your reddits, and facebooks, and twitters, and SaaS-comment fields, and discords. They have the economies of scale to develop viable countermeasures, to hire teams of people to work on the problem full time and maybe at least keep up with the ever evolving bots. I only agree when it comes to the system resources that can keep up with bots. When it comes to fighting spam, these services often do a terrible job because 1) their business model benefits from higher user & engagement numbers and 2) their monopoly status affords them to retain users even if their experience is degraded by the spam, something a small site often won't be able to do.
- kazinator 4y agoIn the 1980's, we kept anklebyters off dial-up BBSses with a simple technique: voice validation. To join the forum, you had to fill an application first, which included your real name and phone number. The sysop would give you a call for a quick chat, and then grant you access if you didn't seem like a twit. This would be entirely practical for some small-time operator trying to run a forum off residential broadband, while impractical for the reddits, facebooks and twitters.
- s1k3s 4y agoI guess it's a different time and it also depends on who's your target audience. Some people go crazy if you ask for their email address. Phone numbers and calling is a big no-no.
- mnd999 4y agoPhone number is an excellent tracking identifier across services. Even better than email, which is why the data hoarders want it.
- mjevans 4y agoI'm one of those radical militants who refuses to give up any means of direct contact... However for a small scale thing I'd gladly go visit at a face to face meetup to fulfill this type of validation.
- shaburn 4y agoWhat if they came to you. What is the imputed value of that network connection relative to cost...?
- nottorp 4y ago> However for a small scale thing I'd gladly go visit at a face to face meetup to fulfill this type of validation. Even if it were 3 flights totalling 18 hours away? :) Or even just from one coast of the US to another...
- deleted 4y ago[deleted]
- timmaxw 4y agoI wonder if proof-of-work would help. Suppose every form submission requires an expensive calculation, calibrated to take about 1 second on a typical modern computer/smartphone. For human users, this happens in the background, although it makes the website feel slower. But for bots, it dramatically limits how many submissions each botnet host can make to random websites.
- tmikaeld 4y ago"mCaptcha uses SHA256 based proof-of-work(PoW) to rate limit users." https://github.com/mCaptcha/mCaptcha https://github.com/mCaptcha/mCaptcha
- timmaxw 4y agoNice! Yeah, mCaptcha looks like just what I had in mind. I wonder why this approach hasn't been widely adopted?
- tmikaeld 4y agoProbably due to "PoW" being power-hungry, but that's largely false because you only apply PoW here on users that are abusing the system. Allowing abusers to freely abuse would cost even more power than just forcing them to do the work.
- rapnie 4y agomCaptcha is in the process of being adopted in Gitea and Codeberg. See recent Fediverse post from the project account: https://gts.batsense.net/@mcaptcha/statuses/01G9KRBRC8CRC9M3KCVG293HGN https://gts.batsense.net/@mcaptcha/statuses/01G9KRBRC8CRC9M3...
- realaravinth 4y agoThe project is very new, I haven't started promoting yet. The Codeberg development was purely from word of mouth :) disclosure: I'm the author of mCaptcha
- realaravinth 4y ago
- minimalist 4y agoIt is interesting to watch comments about this dance around the topic of barriers to entry. It wasn't exactly easy for the uninitiated to access various internet fora in the early days and with popularity comes the bots born out of the desire to profit for little work at the expense of the community garden. The recent story about VRchat embracing anticheat DRM is another example of this, as its ascending popularity led to more scammers [0]. Does this extend to societies as well? One can think of a membrane that has selective permeability to ideas but resists antisocial actors and concepts. Alexander Bard has talked a lot about social membranics (it's a bit hard to search for). As odious as the web3 charlatanry is, I'm starting to yearn anything that raises the transaction costs for the dumbest bots. I remember reading something about new ideas with distributed moderation at some point--maybe someone can refresh my memory. [0]: https://news.ycombinator.com/item?id=32232974 https://news.ycombinator.com/item?id=32232974
- JohnJamesRambo 4y agoWhat is the reason behind bots spamming marginalia? What’s the motivation? What do they gain? I always wonder about these things.
- onefuncman 4y agoI want to run a honeypot for doing more research on bots and the economics for them, but I get bogged down quickly in the planning stages. I should just start with a vulnerable wordpress site or something.
- prox 4y agoWordpress is perfect for this. The amount of bots trying to get in is insane. Like up to 80 login tries on some days for a small potato website. There are also some vulnerable plugins still out there if you actually want them to hack it.
- SyneRyder 4y agoJust make a site with a Contact page, with a comment form that logs the details of every request (IP address, timestamp, message content, email provided). You'll get plenty of data for research, once the page has been indexed into the database the comment form spammers use. For bonus points, put the contact form at the bottom of every page of your website. A couple of my toy/project websites accidentally became honeypots. Rather than shut down the comment forms, I now have those sites generate summary logfiles that I can upload daily to AbuseIPDB. EDIT: Forgot to mention, also log the Referer field and User-Agent on each request. Very, very useful information for research and detection.
- Avamander 4y agoThere are multiple reasons - negative SEO, positive SEO, malware distribution, paid clicks, advertising and probably others I've forgotten at the moment.
- marginalia_nu 4y agoSimple answer is I don't know, but it appears to be happening to other search engines as well. My best guess is they're assuming it's backed by google, and are attempting to poison its search term suggestions. The queries I've been getting are fairly long and highly specific, often within e-pharma or online casino or similarly sketchy areas.
- fabianhjr 4y agoIt would be simpler to decentralize and implement webs of trust[1] (that locality would also help community-building / social cohesion). Secure Scuttlebutt[1] doesn't have a lack of moderation / spam issue and it is completely decentralized and without monetary fees nor proof-of work. Why can't centralized services do better? [1]: https://ssbc.github.io/scuttlebutt-protocol-guide/#follow-graph https://ssbc.github.io/scuttlebutt-protocol-guide/#follow-gr...
- BiteCode_dev 4y agoIt's not that bad. First, of course, you have cloudflare and recaptcha, which are free and very efficient, as the author say. But even if you don't want to use them (some of my services don't), most bots are very dumb: - require JS, and you lose half of the web ones - silly tricks like hidden input fields in forms that worked in 2000 still work in 2022. Use a bunch of them, and you can yet again halve the bot traffic. - many URL should have impossible to guess paths. E.G: just changing the /admin/ url to a uuid in django or the /wp-admin/ in wordpress, you save so many requests. - bots are usually not tailored to your site, meaning if you require JS, you can actually embed anti-bot measure in the client code and they will work. E.G: exponential backoff + some heavy calculations if too many fast consecutive ajax requests. - fail2ban + a few iptables rules (mitigate syn flood, etc) will help - varnish + redis gets you very far to shave excess dummy traffic It's not great, but it's not an apocalypse. Unless you are under targeted attack. Then it sucks and you die.
- efitz 4y agoAlso, attackers are rarely going to try to guess your URLs - they’re going to find them via Google or Shodan, or, if you’re a good rest citizen, via “/<yourapp>/“
- bryanrasmussen 4y ago>Also, attackers are rarely going to try to guess your URLs - because then the attack becomes DOS as they cycle through dictionaries of words?
- raverbashing 4y agoAny website gets probed for wp-admin.php etc, even if you don't use WP
- BiteCode_dev 4y agoIn fact, if someone is probing for wp-admin, you should insta ban them, no matter the site.
- 4y ago
- Joel_Mckay 4y agoFor small sites, I would just use a simple firewall: 1. whitelist the finite IP ranges for the regional ISPs/country where you do business 2. blacklist the proxy and tor exit nodes 3. blacklist the list of published compromised servers 4. add spamhaus blacklists 5. add fail2ban rules to trip on common server security scans, and unused common service ports 6. publicly reply to those having access issues, and imply they have bad neighbors. This will often take care of 99% of the nuisance traffic, but I still recommend live monitoring traffic regularly. ;)
- philprx 4y agoTor users are often legitimate good internet citizens. A lot of (lucky) us have the luxury to live in real democracies. Some others live in countries that use every single aspect of their private lives (DPI, mass surveillance) to put pressure on them and bend them to the regime's will. In my opinion, Tor and anonymity should not be killed as a result of silly bots.
- Joel_Mckay 4y agoYour opinion is duly noted, and I agree most knowledge should be equally accessible to give everyone a chance to grow. That being said, a commercial site owes nothing to financially irrelevant bandits, sociopaths, or shills. Try it for a week, and then weigh the liability again. ;)
- RL_Quine 4y ago> Tor users are often legitimate good internet citizens. We have had exactly zero traffic from it at any point which was legitimate. Any user who ever showed up with a exit IP ended up being banned eventually, so we just proactively fraud banned anybody who uses one, and anybody that was related to them. There is zero value in allowing anonymizer traffic on your service, and a whole lot to lose.
- Animats 4y agoWhat's hard to do now is host a lightly used but broadly interesting service that doesn't require a login. Although, surprisingly, I host such a service, and while it gets a constant stream of random hits, they're a minor nuisance. Probably because it's just the back end for a web page, and nobody bothers to target it specifically. Random web browsing won't find it, and the API will just return an error if called incorrectly. Even if it is called correctly, it has fair queuing on the service, so hammering on it from a small number of IP addresses won't do much. That did happen once. Someone from a university was making requests at a high rate and not even reading the results. I noticed after a month, and wrote to their department chair, which stopped the problem.
- s1k3s 4y agoYes, this is why I plan to take down my hobby projects. And it's not only bots, real people do it as well. Apparently some people have a passion for screwing up other people's work. Some even email me afterwards asking for money to disclose a bug they found.
- closedloop129 4y ago>What's hard to do now is host a lightly used but broadly interesting service that doesn't require a login. Which other broadly interesting services do exist? The owners of those services could come together and offer a VPN that gets preferred treatment for these services. This could be more precise than https://www.abuseipdb.com/ https://www.abuseipdb.com/.
- Aachen 4y agoSame! I also got like 20 requests every second from a university IP. I tried a few things to make it error out, like returning 404, but no dice. In my case, it was my own fault though, a page with a few lines of JS to periodically check for updates got into a crazy state (I never found out how) and they didn't notice because it was a remote desktop system where they left the page open. Went on for months but didn't impact my service (I just noticed it in access logs while looking for something else) so I left it and remembered again a few months later, then it was gone.
- david_draco 4y agoHave a "CAPTCHA" that gives the IP reputation for some time (cookie+IP=key), but instead of a CAPTCHA make the web page / browser solve and submit a BOINC task from a randomly picked science project. No user interaction needed, it has the benefits of "paying by computation" of cryptocurrencies without the tracing, and if bots solve the problem efficiently, it's good for science.
- rapnie 4y agoThat is a nice idea. So bit similar to mCaptcha [0] that uses PoW algorithm, mentioned in other comment [1] in the thread. [0] https://mcaptcha.org/ https://mcaptcha.org/ [1] https://news.ycombinator.com/item?id=32339902 https://news.ycombinator.com/item?id=32339902
- zakki 4y agoCan we make a bot to mine a cryptocurrency?
- GTP 4y agoIt's called miner and you can already install it on your pc.
- GTP 4y agoBut solving a BOINC task requires too much time while the average user rightfully expects a webpage to load within 5 seconds or so
- david_draco 4y agoUsually you have a landing page, and then you enter stuff there, and finally you get output. That gives you about 1 minute before returning first results. If the user is faster, you can show something like cloudflare does when you visit through Tor. On later submissions you can reuse the reputation from the cookie.
- GTP 4y ago
- jb1991 4y ago> They're a major part in killing off web forums, I’ve noticed that a lot of old popular forums disappeared in recent years, but I didn’t realize it was possibly due to bots. Why is that? I assumed that the admins just got tired of running them and moderating them.
- marginalia_nu 4y agoIt's more complicated than just bots, competition from Reddit is another factor, but bot traffic were certainly a significant part of the problem, both in terms of the constant drive-by exploits and ceaseless comment spam drove up the amount of work needed to operate a forum as a hobby to basically a full time job. With waning visitor numbers, it simply became untenable.
- golergka 4y ago> If Marginalia Search didn't use Cloudflare, it couldn't serve traffic. There has been upwards of 15 queries per second from bots. 15 RPS is very far from an apocalypse.
- deleted 4y ago[deleted]
- Avamander 4y agoIt's bad if it's your dead-average Wordpress site that has 10 PHP workers, each page load being >1s. Easy DoS.
- Aachen 4y agoYeah but WordPress is an extreme example. Every time a WP blog is posted to HN without a static-page-ifier (caching layer that basically turns the dynamic pages into static ones), it dies within minutes. Normal software doesn't seem to have that problem. I traced it once, and I got to admit there was not an obvious bottleneck (this was 2015 or so). Just millions upon millions of calls into deeper and deeper layers for things like translations or themes. Wrapping mysql_query in a function that caches the result (to avoid doing identical queries) helped a few % I think, but aside from major changes like patching out the entire translation system for single-language sites, I didn't spot an obvious way to fix it. You'd need to spend a lot of time to optimize away the complexity that grew from suiting a million different needs, contributed by thousands of people across many years.
- marginalia_nu 4y agoIt is if you're hosting an internet search engine on a PC.
- deleted 4y ago[deleted]
- golergka 4y agoWhy would you do such a thing in the first place?
- SmileyJames 4y agoI thought a plan for spam had solved this one? http://www.paulgraham.com/spam.html http://www.paulgraham.com/spam.html Has NLP progressed rendering Paul's plan a failure? Am I a bot? How about you? Does it matter if I make valuable contributions?
- greazy 4y agoSpam and bots eating traffic are two different things.
- viraptor 4y ago> If Marginalia Search didn't use Cloudflare, it couldn't serve traffic. Cloudflare is not the only CDN/protection. It's the most popular and the most evil one. You have a choice.
- matkoniecz 4y agoWhat alternatives you recommend?
- viraptor 4y agoIt depends on your audience and regions you're most interested in. But if you're aiming for the EU, gcore labs may be interesting. Akamai is not bad, but a bit enterprisey - I don't think they even had an official api the last time I used them?
- randunel 4y agoNone of those are free, though.
- viraptor 4y agoNo, but they also don't actively help protect pages organising SWATing. It's your choice who to do business with.
- Aachen 4y agoIf you're not paying, what's the product they're selling?
- RL_Quine 4y agoTheir paid one when you go over the limits. Not everything has to be black and white and reduced down to a single, oft repeated catch phrase like that.
- RL_Quine 4y ago
- nixcraft 4y agoI run a popular blog and confirm that spam is a massive issue. I am trying to keep the independent web alive with an old-school commenting system because it helps readers and myself improve outdated posts. My domain is over 20+ years old and attracts all sorts of threats, including monthly DDoS and daily spam. Using Cloudflare solved all of these problems. Next, you need to add firewall rules inside Cloudflare WAF to trigger a captcha for /path/to/blog/wp-comments-post.php. That will not get rid of human spam tho. For that, you need to use another filtering service called Akismet.
- toastal 4y agoPutting everything 'behind Cloudflare' isn't a panacea. By merely living outside the West, I'm getting Geo blocked from 'normal' news sites and constantly having to solve hCAPTCHAs to solve riddles for some AI algo without compensation. It's such a burden and I find myself giving up pretty often. GeoIP blocking is what prevented me from getting my voter information out of my last domicile. Running everything through Cloudflare or similar also contributes to the concept of letting the internet be centralized around a few choke points that can hurt free speech (both the good and bad kind) and when they go out (which did recently) a large swath of the internet comes with it.
- jks 4y agoDoes Cloudflare's "Privacy Pass" browser plugin help at all? It's advertised as reducing the number of hCaptchas you need to solve by a factor of 30, but I rarely see hCaptchas anywhere on my connection so I can't really evaluate myself.
- nixcraft 4y agoI agree with you. But, what solution do you propose for independent solo developers or people who wish to run a blog instead of using FB, Twitter and co to create content? Cloudflare may not be perfect, but it prevented me from shutting down my solo operation without putting a massive cost burden on me. When the first time DDoS hit, I had to beg one of those large cloud companies to reduce bandwidth costs. It took them forever to forgive that abuse and price, which was not my fault, and I was given a strong warning not to repeat such an issue again. There is no easy solution to this problem. At least with Cloudflare, people like me can stay online, but it does cause a problem for a bad IP reputation. TL;DR: I won't expose any of my projects or API directly these days due to spam, ddos and other abuse.
- Avamander 4y agoIt's annoying for sure. I deal with abuse at a large scale. I'd recommend: - Rate-limit everything, absolutely everything. Set sane limits. - Rate-limit POST requests harder. Preferably dynamically based on geoip. - Rate-limit login and comment POST requests even harder. Ban IPs that exceed the amount. - Require TLS. Drop TLSv1.0 and TLSv1.1. Bots certainly break. - Require SNI. Do not reply without SNI (nginx has 444 return code for that). Ban IP's on first hit that connect without. There's no legitimate use and you'll also disappear from places like Shodan. - If you can, require HTTP/2.0. Bots break. - Ban IP's listed on StopForumSpam, ban destination e-mail addresses listed there. If possible also contribute back to SFS and AbuseIPDB. - Collect JA3 hashes, figure out malicious ones, ban IPs that use those hashes. This blocks a lot of shit trivially because targeting tools instead of behaviour is accurate.
- gkbrk 4y ago> If you can, require HTTP/2.0. Bots break. Non-bots break as well. I have Firefox configured to use HTTP/1.1 only. No reason to chase Google's standard-of-the-day, HTTP/1.1 has worked for ages and it will continue to do so for the foreseeable future.
- sirshmooey 4y agoGenuinely curious, why disable HTTP2? Your web browsing must be awfully slow sans multiplexing.
- gkbrk 4y ago> why disable HTTP2 Because it adds nothing to improve my browsing experience, and reducing the number of protocols supported by my browser from 3 to 1 also reduces the attack surface. > Your web browsing must be awfully slow sans multiplexing. And yet it's not slowed down at all. How many different resources must a web page use before it feels slow on a connection pool of keep-alive TCP sockets? Maybe people visit some wild experimental web pages with hundreds of blocking <src> tags that are not bundled/minified? Either way, my experience is it doesn't slow anything down when I use both websites (forums, resources, youtube, social media) and web apps (banking, maps, food delivery etc).
- 12907835202 4y agoFor my forum with 500k users a month I just added a registration captcha related to my niche. E.g. for a Dark Souls forum it would say "what game is this forum about?" And if you got it wrong the validation would include "tip it's just two words Drk S*ls". This reduced spam by over 99% and didn't annoy people with recaptcha. If someone was unable to get past that captcha (it still happens I have logs!) I figured they were probably not that valuable a contributor anyway. If someone wanted to target my site directly they could but hasn't happened so far.
- d3nj4l 4y agoA niche dark souls forum sounds interesting, any chance I could get a link?
- google234123 4y agoYou misread the post. That was just an example. A niche dark souls forum wouldn't have 500k users lol.
- Lex-2008 4y agore: someone was unable to get past that captcha - this reminded me a story I heard back in ICQ times about some human who couldn't pass anti-bot question: "What planet do we live on?"
- gilrain 4y agoFair enough… one can only speak for oneself, after all.
- Suzuran 4y agoI remember a friend's con-group forums who had an issue along these lines - the anti-bot question was "What is the brightest thing in the sky at noon?" the expected answer was "the sun", but some guy got stuck because he was answering "Sol". Since they had an IRC channel the issue was relatively quickly resolved, but it was an in-joke for some time.
- IfOnlyYouKnew 4y agoSo there’s one service keeping this search engine online, and it’s probably doing it for free, and the author can’t even think of a better way to do it. Yet Cloudflare still gets two paragraphs of complaints in the face? Because the author wants to “own” something instead of “renting”?
- marginalia_nu 4y agoI'm doing it for free because I don't want this to be a commercial service. I get that HN is startup city, but I'm not running a startup, it's just a hobby.
- lloydatkinson 4y agoI found that my netlify site attracts a lot of spam specifically from the same spammer/group. The messages always start with some variation of “Hi my name is Eric”. Netlify seem to not really care after reporting it on their support forum. The spammer disables JS so no client side protection works. I’ve recently decided to (unfortunately) break the ability for JS disabled browsers to be able to submit the contact form. The form elements attributes are all wrong meaning the form won’t submit correctly. Instead, some JS on page load sets the attributes to the correct values. I will wait a while and see if this solves it. While netlify does correctly mark all this as spam the fact is that legitimate messages sometimes can slip past these with false positives. So I have to check the vast amount of spam often.
- 1vuio0pswjnm7 4y ago"The rest are forced to build web services with no interactivity, or seek shelter behind something like Cloudflare, which discriminates against specific browser configurations and uses IP reputation to selectively filter traffic." Interactivity is not a must-have. The world's first general purpose computer, ENIAC, was not built for "interactivity". It was built to calculate ballistic trajectories, which were otherwise calculated manually. Computers exist to allow automation, to reduce manual labour.^1 "Tech" companies need interactivity to support collection of data about www users and paid services related to programmatic online advertising. Generally, users do not need interactivity. Generally, users do not need to spend excessive quantities of time "interacting" with networked computers. As a user, I want _non-interactive_ www services, whether it is data/information retrieval or e-commerce. I want to use more automation, not less. Automation is not reserved for those providing "services". It also should be available to those using them. Provide bulk data access. Let others mirror it. Take advantage of "open datasets" hosting if necessary. For example, Common Crawl is hosted for free with Amazon. Upload the data to internet Archive. "The API gateway is another stab at this, you get to choose from either a public API with a common rate limit, or revealing your identity with an API key (and sacrificing anonymity)." Publish the rate limit for the public API. Do not make users guess. Do not require "sign-in" to use an API to retrieve public data. 1. Some folks consider having to "interact" with a computer as labour, not fun.
- lifeisstillgood 4y ago>>> Automation is not reserved for those providing "services". It also should be available to those using them. Yes ! I call this software literacy. And yes - no matter how cool the JS on a major site, the fact that the sites goals are to keep me there and clicking and my goals are to get what I want with minimal action are in conflict. I would suggest that bots are actually not a problem. For most things I would like a bot acting for me. Telling me as and when that I need to visit the dentist, who has slots free next weds and friday. Friday is best because I am also WFH that day. The bot apocalypse is only one because we are trying to make a "web for humans" when actually a "web for bots, and a bot for a human" is a much better idea :/)
- 4y ago
- JimWestergren 4y agoI am running a website builder with > 20K sites. I use open contact forms without captcha. What worked for me is to use a one line javascript that places current timestamp in a hidden input field that is default 0. Then I check on the backend and if the value is either 0 or time to fill out and send the form is less than 4 seconds I block as spam. This blocks more than 99% of spam and also takes care of most human copy paste spam as well.
- naillo 4y agoI like this solution because spammers are unlikely to try to get around it. A delay eats into their time budget and they can't introduce a human-like waiting time on every site they try to spam, better to just move on to find cheaper targets.
- walls 4y agoYou could just decrease the timestamp instead of actually waiting.
- naillo 4y agoI meant for general spammers who goes after tons of sites mostly blind. I agree it would not help for a targeted attack.
- JZerf 4y agoI'm already using this timestamp technique on my website and so far no bot operator has bothered trying to work around this. However even if some bot operator were to specifically target a website using this technique and try to decrease the timestamp, I believe you could still force a bot to wait by just changing the website to use something like a cryptographic nonce that includes a timestamp instead of just a simple timestamp that can be understood easily.
- robalni 4y agoIf you don't want to require users to run javascript you should be able to make the server generate the timestamp.
- GeckoEidechse 4y agoI wonder if a general solution could be to make the visit more computationally demanding to the visitor than to the host, e.g. some form of proof-of-work. I guess captchas already do that in some sense but they require the humans to do the work. Now the author above has stated they dislike the crypto route and I agree that the whole web3 idea is bs but what if in the case that spam of some form is detected by the server, it requires the visitor to show some proof-of-work and combine that with the "mining crypto in JS instead of ads" craze. That way the bot would need to put work in which would slow it down and at the same time it would pay for its own visit. No ofc no spam detection system is perfect and it would also hit human users but in their case it would be just a wait a few more seconds longer for page to load kinda case.
- Bloating 4y agopost this above, but there is a pre-bitcoin whitepaper suggesting just this approach to solving email spam
- lifeisstillgood 4y agoI would suggest that bots are actually not the underlying problem. For most things I would like a bot acting for me. Telling me as and when that I need to visit the dentist, who has slots free next weds and friday. Friday is best because I am also WFH that day. The bot apocalypse is only one because we are trying to make a "web for humans" when actually a "web for bots, and a bot for a human" is a much better idea :/) We need to redesign a web based on APIs, certificates, rate limits etc. And stop having "engagement" as a goal, and have "getting things done" as a goal Edit: mucked up formatting
- danrl 4y agoOff topic: Great to see more Gemini/Web dual hosted sites.
- jgalt212 4y ago> I can't afford to operate a datacenter to cater to traffic that isn't even human. This spam traffic is all from botnets with IPs all over the world. In our experience (we don't have a forum), almost all of our bot traffic has been SEO spiders (or claiming to be so).
- SyneRyder 4y agoReally glad to see someone finally talking about this. Does anyone know what's going on with that "Duke de Montosier" spam botnet? It accounts for more than half of the botspam attacks on my sites, and I can't find anyone talking about it online anywhere, except one tweet dating back to mid-2021. It's identifiable by several short phrases that it posts: Duke de Montosier for Countess Louise of Savoy Testaru. Best known And cryptic short posts that can assemble into creepy sequences: Europe, and in Ancient Russia Century to a kind of destruction: Western Europe also formed and was erased, and on cleaned only a few survived number of surviving European 55 thousand Greek, 30 thousand Armenian Many of the IPs involved seemed to be in Russia, China and Hong Kong, though they're coming from all over (eg European & US VPNs, Tor exit nodes). From tracking the IPs on AbuseIPDB, the weird spam posts seem to be just one layer, while behind the scenes it also attempts SMTP Auth and IMAP attacks on the server. I'm eager to know more if anyone knows, and especially if anyone is trying to shut this thing down. But I can't find anyone even talking about it. (Maybe there's a reason for that?)
- prepend 4y agoI assume it’s time travelers trying to post enough so their message persists.
- marginalia_nu 4y agoHow very numbers station of them. I've seen it suggested that botnets use comment fields for command and control, maybe something like that?
- SyneRyder 4y agoMy theory for the phrases above is that they're a "unique seed" used to identify sites that are easily compromised. Do a web search, find a website filled with "Duke de Montosier" comments - bingo, you've identified an easy website to target with your backlink comment spam. Or, more maliciously, a website that is easy to thoroughly compromise with vulnerabilities. But that's just my current theory. Here's the one tweet I found in Swedish about the comment spam botnet, and it dates back to February 2021. She's the only person I could find who has mentioned it in public. Or maybe my search skills are failing me. https://twitter.com/aureliagu/status/1357368329573400578 https://twitter.com/aureliagu/status/1357368329573400578
- Aachen 4y ago> There has been upwards of 15 queries per second from bots. There is just no way to deal with that sort of traffic, barely even to reject it. If the queries are not a megabit each, you're doing way too much processing before applying rate limiting. Rejecting traffic ought not to take more than 1-2 milliseconds, even if you need to look up an api key or IP address in the database. I, too, host services on a residential connection: 50 mbps shared with other users. My domains must host hundreds of separate scripts, a few of which have a database attached (I can think of six off the top of my head, but there's over a hundred databases in mariadb so I'm sure there's more that I've forgotten about). This is a ten-year-old laptop with a regular "apt install mariadb", no special configs. Yes, most traffic is bots, and yes sometimes they submit more than 1 q/s. But it comes nowhere near to exhausting resources to a noticeable extent. Load average is about 0.15, main peaks come from my own cronjobs. If you're having this much trouble rejecting traffic, you might want to spend some time looking at bottlenecks. You'll also notice the bots knock it off if it's unsuccessful.
- marginalia_nu 4y agoThat's queries per second (as searches), not requests per second. You know how blogs sometimes can't handle being on the hacker news front page from all the traffic? Well my search engine has survived that. That was 1-2 QPS. Unmitigated bot-traffic is roughly 10x the traffic of a hacker news death hug, as a sustained load.
- shp0ngle 4y agoThe actual blogpost aside: the margianalia search is the first of these “alternative search engines” that I actually like. Most of those has been either “worse google” or “utter trash”… this one returns some interesting results for some queries I have tried.
- AviationAtom 4y agoIt's funny the author mentions Facebook and Twitter, because the bot spam on both is quite apparent. The spam on the former has risen greatly, seemingly mostly from India and parts of Africa. Scam air duct cleaning posts, posts about hacked account recovery, and random other crap. It really degrades the experience of the Internet, IMHO.
- FrenchDevRemote 4y agoDoes anyone know how google/linkedin manage to block bots who are using SSO? Trying to login to a linkedin account using a google account from an automated browser(like puppeteer+puppeteer-stealth or fakebrowser), will open a white empty window instead of the normal google login window, I could be a limitation of those libraries, but I doubt it, smells like something they detect, maybe looking into might lead some interesting insights on how to limit modern bots.
- rrwo 4y agoI run a website for a small company. The site has been around since the mid-1990s, and bots are a minor annoyance, but not a problem. We also use some simple heuristics to reject obvious bot traffic. One of the simplest is to have a form field that is hidden via CSS. Humans don't see it and it stays blank. Bots fill it in. Bots tend to fill in every form fields with random garbage, even checkboxes. Validating checkboxes rather than checking they have a value is another good way to detect bots. Many bots have a hard time with CSRF tokens in hidden fields. Many bots also don't handle session cookies properly. If someone submits a registration form without an existing session, we reject it. (So we don't get as far as checking the CSRF token.) After a certain number of failed attempts to register or login, we block the IP for a period of time.
- SahAssar 4y ago> There has been upwards of 15 queries per second from bots. There is just no way to deal with that sort of traffic, barely even to reject it. I don't really understand, is that a lot? 15qps does not sound like a lot, especially for a blocking/rejection function.
- marginalia_nu 4y agoIt's 15 search queries per second, not requests per second. RPS is usually 10-20x higher.
- SahAssar 4y agoBut you said "barely even to reject it", rejecting 15 QPS should not be heavy on any resource, right? Or is the actual problem identifying the bot traffic?
- SavageBeast 4y agoI get paid over 92 Dollars per hour working from home with 2 kids at home. i never thought i'd be able to do it but my best friend earns over 15k a month doing this and she convinced me to try. the potential with this is endless... Simply go to the BELOW LINK and start your work.. EDIT: bad joke but maybe someone will get a chuckle.
- Kiro 4y agoCoin Hive was an interesting solution before it became synonymous with crypto jacking. In order to post a comment you had to lend your CPU to mine for X seconds. The only true anonymous and frictionless micropayment system I've seen.
- figmaheart255 4y agoWhile clearly bot spam is on the rise, we need to be very careful on how we choose to deal with it. Cloudflare has already introduced "proof-of-Apple" [1], where proven Apple devices get special treatment, bypassing captchas. Later we might see websites that are only accessible via Google, Microsoft, or Apple devices. If we continue down this path, we'll end up with a social credit system ruled by big tech. [1]: https://news.ycombinator.com/item?id=31751203 https://news.ycombinator.com/item?id=31751203
- kube-system 4y agoWe basically already have "social credit" systems, we just call them anti-fraud/anti-spam/reputation scores.
- js4ever 4y ago"There has been upwards of 15 queries per second from bots. There is just no way to deal with that sort of traffic, barely even to reject it." What??? My phone can serve that easily, any modern server can handle 50-250 rps
- marginalia_nu 4y agoThis is queries per second (as in I run a search engine), not requests per second.
- julianlam 4y agoI disagree with TFA's take on dealing with spam — giving up! For our app, we don't deal with spam in any novel way. We use honey pot, SFS, and Akismet. However, by far the easiest way to stop spammers is a post queue. Lots of spammers will just create a burner account, fire off their spam, and start over. Given no actual reputation, give them the trust they deserve — none. The other factor is building out a fast backend. Besides benefiting your own users, it also means Googlebot or Ahrefsbot won't absolutely cripple your site when they come knocking. Sometimes that is doable, sometimes not.
- marginalia_nu 4y ago(Author) I'm running a search engine though. Do you propose I require users to register an account, and then not allow them to search? I think my backend is plenty fast given it's hosted on a PC off domestic broadband. Most searches complete sub-100ms.
- julianlam 4y agoHey, thanks for the reply! Specific scenarios require creative solutions. For a search engine, how do you differentiate between robots and legitimate users? It seems a rate limiting step is likely the best solution. If query rate from the same IP exceeds a threshold, throttle them creatively. +100ms the next time, +250ms the next, etc. The upside is these bots will adjust their strategies to hit your site slower, which is the whole point, isn't it. If they spread requests across IPs, perhaps try fingerprinting. I'm not sure how effective that is on the backend though.
- lizardactivist 4y agoEnd game: everything runs on US-owned services, and all users need to identify to be allowed to even raise a finger, so that "bad actors" can be kept out. All while we blame Russia and China, and say that their spambots and evil actions forced us to do this.
- 16amxn16 4y agoThis actually makes sense. Or, at the very least, it wouldn't surprise me. Another end game: some sort of ID is required to use anything (that ID being a local phone number, which can be tracked down to you).
- djohnston 4y agoAre you suggesting that the U.S. produces a proportionally similar volume of spam traffic as Russia, China, India, Vietnam?
- boredumb 4y agoI get a ton of spam from my contact me pages even with a captcha in place, i've been experimenting with loading an initial dummy form and replacing it within a few seconds of loading to the real deal which seems to have cut down on bots submitting stuff. Rate limit everything you can and use a captcha where acceptable, there are also a load of public IP and email blacklists that you can use to run a quick check. Working in a field where there is a large amount of bots and incentive to abuse we invest quite a bit of time and money in fraudulent traffic detection using a cornocopia of different services in tangent and at the end of the day we still see a small percentage of traffic getting through that is fantastically human like. With that out of the way, I've been engulfed in AI and GPT3 functionality lately and I thought this post was going to be doomsaying the coming apocalypse of bot spam, because the level of human like quality coming from the AI is going (already has) to make deciphering human vs bot traffic/posts/emails/comments nearly impossible. It will be fun here soon when we see forums entirely dedicated to bots conversing and arguing with each other outside of reddit.
- ComputerCat 4y agoSame! The captcha doesn't seem to be able to slow down the bots. Inbox is still getting flooded with spam.
- deleted 4y ago[deleted]
- djohnston 4y agoI work in this space at a company you've heard of - even at our scale and with our resources the proportionally larger attack incentives mean we are constantly firefighting. > The other alternatives all suck to the extent of my knowledge, they're either prohibitively convoluted, or web3 cryptocurrency micro-transaction nonsense that while sure it would work, also monetizes every single interaction in a way that is more dystopian than the actual skull-crushing robot apocalypse. I understand the drawback here but I would like to see monetized transactions employed as a defense layer a little more before we make a final decision. It is undemocratic, to be sure, but maybe for those of us who can afford it it's still better than the cesspool we currently sift through on every major platform. Anyone aware of any platforms taking this approach? Maybe the fediverse will help - by fragmenting networks attackers may have less incentive to attack a particular one.
- reaperducer 4y agoAnyone aware of any major platforms taking this approach? The Postal Service? Sure, there's junk mail, but imagine how much junk mail there would be if it were delivered for free. It wasn't until phone calls became so cheap as to be "unlimited" that we ended up flooded with billions of junk calls. Microtransactions (non-crypto, thankyouverymuch) would solve a certain number of today's problems.
- marginalia_nu 4y agoI do think it would help, like even if a transaction cost 0.05c, it would add up very quickly for a bot operator but stay cheap for everyone else. But I think the problem is it would inevitably introduce the need for a middle man, shaving 0.01c off that 0.05c, with a dubious incentive to increase the amount of money changing hands as much as possible. What you've invented at that point is basically Cloudflare with worse incentives. You get either that, or yucky defi web3 crap.
- djohnston 4y agoYes for sure, I have thought about making an email "stamp" web-3 service that would implement this. I even wanted to make some fun "pony express" animations whenever a letter was arriving to your inbox.
- Xeoncross 4y agoI have considered skipping the regular fingerprinting, geolocation, captcha, hashcash, email verification, payment required, etc... mitigations and instead requiring people to drop into a public chat room (or pm/chat the support team) to have their account activated. The number of languages supported would be small to match whoever helped moderate this, but it would at least require speaking to someone. A PM thread or live chat would be an instant way to find out if someone can string two sentences together and might be worth allowing into the site. You could even have them create an account and solve a single captcha prior to getting access to the chat. It's not perfect by any stretch, but might be worth exploring having humans-verify-humans.
- z3t4 4y agoMy trick is to have one field that should always be blank and one field that should always have an value, this stops all automated bots. No "CAPTCHA" needed.
- ltr_ 4y agotangential : Two weeks ago (and for a while) our country's twitter-sphere (Chile) was completely and obviously dominated by bots, they were starting and inflating trending topics with absurd lies, spreading fear and chaos in favor of "Rechazo" (the option against our new constitution in the next ballot) or echo chambers for republican and extreme right associated politicians.. What happened? a self organized group[1] started to do data analysis of the trending topics and delivering the results to the people showing who was behind the campaigns and synthetic likes, after this, prominent and public figures from that sector started to cut funding for bot networks (because of the public shaming and media attention they were receiving) and is so pathetic now ,they can't even get more than 100 likes and often the most popular response is a refutation or the very same analysis showing the bot network working with substantially more organic likes. I think is a very interesting phenomenon to watch. Note that this lies/fear/chaos campaign is transversal, from rural AM radio to tiktok, but is not working at all. People is very aware of these campaigns and knows how to defend against. Truth is stronger than money. - [1] https://twitter.com/BotCheckerCL https://twitter.com/BotCheckerCL
- EGreg 4y agoI will reiterate what I had been saying on HN for years: 1) The problem is centralization. Yes DNS is federated but there is a central registry. This means anyone can spam you@yourdomain.com or visit your web server listening for HTTP connections at www.domain.com 2) DNS is a glorified search engine. Human readable domain names are only needed for dictating a domain name (and listeners often make mistakes anyway). They only map to a small fraction of URLs) namely the ones with “/“ path name. For most others, the human readability adds little benefit. 3) Start using URIs that are not human readable. The titles, favicons and other metadata of resources should simply be cached, and displayed to the user in their own bookmarks, search engines or whatever. For Javascript environments, variables can easily hold non human readable URIs. Also QR codes can resolve to non human readable URIs. 4) There may be some cookie policy for third party hostnames etc. but just make them non human readable also. 5) We should have DHT or other decentralized systems for routing, and here is the key… in this system, you need a capability issued by the website / mailbox owner in order for your message to be routed to them. If the capability is compromised and used to get a ton of SPAM, they simply revoke that specific capability (key). For HTTP websites you can already implement it on your side by signing the keys / capabilities ie session cookie balues with an HMAC, and there is no need to even do network I/O to verify them, you can upload the whitelist to the edges and check them there easily. But going further, for new routing protocols, IP addresses should be removed after the first hop in the DHT, because the global routing system will send traffic there otherwise. See how SAFE network does it. 6) I don’t need a “real names policy” or “blue checkmark”. I can know who “The Real Bill Gates (TM)” is through some verified claims by Twitter or someone else. Just because I have the email billgates@microsoft.com doesnt mean I should be able to email him. There can be many Bill Gates. The names are just verified claims by some third party. Here on HN we dont have names or photos, and it works just fine. 7) Most of the celebrity culture, papparazzi, Elon Musk and Donald Trump moving markets and tweeting at 5am to 5 million people at once, are problems of centralization. Both a 1 to many megaphone and a many to 1 inbox. Citizens United is just a symptom of the problem. I have spoken about this (privately owning access to an audience) with Noam Chomsky in an interview I did a year ago: https://community.qbix.com/t/freedom-of-speech-and-capitalism-in-2021-interview-with-noam-chomsky-political-commentator/158 https://community.qbix.com/t/freedom-of-speech-and-capitalis... Fox News (Rupert Murdoch), CNN (Ted Turner), Twitter (Elon or Jack), Facebook (Zuck) are controlled by only a few people. Channels on youtube, telegram, podcasts etc are controlled by a few people. This leads to divisions in society, as outrage clickbait rises to the top. Nonprofit models based on collaboration like Wikipedia, Wikinews, Open Source and Science produce far more balanced and benign information for the public. In short we need alternatives to celebrity culture, DNS and other systems that centralize decision making in the hands of a few, or create firehoses and megaphones. Neither the celebrity nor the public actually enjoy the results.
- Test0129 4y agoNot totally unrelated but I had to turn off email alerts and come up with a way to summarize things because Fail2Ban and other alert systems were hit quite literally every 15 seconds with port scans/attempted entries on SSH and other ports. Reporting the abuse to ARIN/ICANN didn't help because almost a full 95% of the traffic originated from China, and 90% of the remaining 5% was Russia. Of the remaining they were zombies inside of America, typically on digital ocean, and I was able to get those handled quickly and efficiently. When I had a simple (secure) login system hosted on HTTPS it was getting hit hard enough my VPS ISP was sending emails to figure out a way to stop it. There are literally 3 people that even know of the existence of these services. It is actually nuts just how much bot spam there is.
- unixbane 4y ago>spam captchas were designed to solve this (and only this, as opposed to requiring them to merely view content like modern ignorant web devs like to do [yes i know some web devs now require it to be able to make sure the people they're datamining are real, but this is a new practice from this year basically]) public services should be implemented by decentralized p2p. static content is solved by ipfs, freenet, etc. dynamic content perhaps can only be solved with smart contracts, which would be less bad than cloudflare if they weren't expensive, as they still provide protocol conformance (unlike cloudflare that requires you to have your packets look like a big 4 browser), anonymity (yeah, pseudonyms, you can still make one per query), etc. without smart contracts many interactive applications are still possible > The other alternatives all suck to the extent of my knowledge, they're either prohibitively convoluted, or web3 cryptocurrency micro-transaction nonsense that while sure it would work, also monetizes every single interaction in a way that is more dystopian than the actual skull-crushing robot apocalypse. centralized web hosting is and always was unsustainable and this is the reason most web content is commercial garbage, and the problem will only get worse. my concern was always what kind of garbage boomer protocol will become the new standard. i sure as hell dont want something that looks like email, web, or UN*X.
- goatcode 4y ago>large resources causing bot spam >large resources are the solution To those who have been recently pondering the history of antivirus companies of the 90s and 00s, and suspiciously wondering how they were always able to so quickly come up with definitions for the newest infections, this all feels so familiar. What a sad world we live in, sometimes.
- runeks 4y ago> The other alternatives all suck to the extent of my knowledge, they're either prohibitively convoluted, or web3 cryptocurrency micro-transaction nonsense that while sure it would work, also monetizes every single interaction in a way that is more dystopian than the actual skull-crushing robot apocalypse. Payment per request is the long term solution, but I completely disagree that it’s in any way dystopian. The trick is to set the fee so low that humans, who make few requests, aren’t really affected while bots, who make a large number of requests, become unprofitable. It’s exactly the same solution as email spam: at a hundredth of a USD cent per email, spam emails would no longer be profitable while regular consumers would spend 10 cents per year (assuming they send 3 emails per day).
- SimplGy 4y agoSo you actually mentioned web3, but not the technology from it that is most interesting to me for this problem. I’m hopeful that “decentralized identity” is a solution here. In theory this would let people cryptographically prove whatever small fact about themselves they would like to share with a service provider like marginalia, without paying you, or sharing their identity, or anything else. Eg: “I am a human”, or “I have written less that 100k words on the Internet” or “my comments have an upvote average above zero”