4 ms·
I recently conducted an experiment - I removed client side CAPTCHAs from a form that had reCAPTCHA V2 and a ton of spam was getting through and instead sent the
by dimmke 3y ago
I recently conducted an experiment - I removed client side CAPTCHAs from a form that had reCAPTCHA V2 and a ton of spam was getting through and instead sent the content to Akismet for scanning. It cut the spam getting through to 0.
It made me think, are client side CAPTCHAs really worth it? They add so much friction (and page weight - reCAPTCHA v3 adds several hundreds KBs) to the experience (especially when you have to solve puzzles or identify objects) and are gamed heavily. I know these get used for more than form submissions, to stop bot sign ups etc…
I feel like it’d be just as/more effective to use other heuristics on the backend: IP Address, blacklisting certain email domains, requiring email validation or phone validation, scanning logs, analyzing content submitted through forms
- kccqzy 3y agoThen you'd just give visitors of your websites no recourse and no information whatsoever on how to fix the problem. The benefit of client-side CAPTCHA is that humans at least can pass it and fix the problem even if something they don't control (such as their IP address having bad reputation due to shitty ISP) is causing problems. As a website operator it's easy to look at the spam that is getting through and be happy that's it's zero. But do you get any idea how many actual humans that you have incorrectly rejected? You don't have that data and it's really easy to screw up there. Of course if your website is small nobody cares. If you are bigger like Stripe you simply get bad publicity on HN. People on HN love to hate on mysterious bans and blocks just because they do something slightly unusual and your backend-only analysis flags them as suspicious. Abuse fighting is hard.
- jsnell 3y agoAnd just to be clear, it doesn't need to be either captchas or doing heuristic abuse detection on the backend. In the ideal case you're making a decision in the backend using all these heuristics and signals, but the outcome is not binary. Instead the outcome of the abuse detection logic is to choose from a range of options like blocking the request entirely, allowing it through, using a captcha, or if you're sophisticated enough doing other kinds of challenges. But proof of work has basically no value as am abuse challenge even in this kind of setup, the economics just can't work.
- deleted 3y ago[deleted]
- dimmke 3y ago>Then you'd just give visitors of your websites no recourse and no information whatsoever on how to fix the problem. This is a weird assumption. What's preventing a backend system from saying "Hey, we think you're a bot. Here's an alternative way to contact us." You obviously don't want to give away enough to help bot developers get through your system, but that's not the same as no resource and no information. >But do you get any idea how many actual humans that you have incorrectly rejected? Yes - like I said in my other comment, this new system actually logs all submissions. It just puts the ones it identifies as spam into a separate folder. Akismet also has the ability to mark things as false positives or false negatives. I think that automated form submissions are very context specific. So, the example I wrote about is for a marketing site, and it's a business that primarily targets other businesses. Most of the spam it gets is for scummy SAAS software, SEO optimization, etc... But my personal website has a very simple subscribe by email form. There were definitely a few spam submissions - someone just blasting out an email address and signing it up to whatever form would accept it. When I implemented double opt in - gone entirely. My larger point was that as an industry, we seem to have just capitulated to client side CAPTCHAs. And it sucks. It's one of the many shitty things about the modern web. But I think it's become just an assumption that it's needed, and we haven't reexamined that assumption in a while. I think it'd almost be better for there to be something could spin up in a container that has a base machine learning model, but can "learn" as you manually indicate messages etc... and then you can also choose a threshold based off your comfort level.
- kccqzy 3y ago> This is a weird assumption. What's preventing a backend system from saying "Hey, we think you're a bot. Here's an alternative way to contact us." Not a weird assumption, but a necessary assumption based on considerations of scale. A small-scale website that doesn't receive too much spam attempt can manually classify spam by human agents. A medium-scale website can have CAPTCHA to let through some visitors and the rest goes to human verification. You appear to be in this bucket. When the scale is huge, no other alternative way to contact exists. CAPTCHA becomes your only tool. In other words, CAPTCHA is only necessary because of scale; what do you think the first A stands for? But because of scale, alternate ways stop working.
- hedora 3y agoClient side captchas are obfuscated code, so the bar of “the user can debug the problem and fix it themselves” is pretty high. Also, reCaptcha definitely engages in both hell-banning and allows incorrect answers to pass the test. I assume the logic for those things is mostly server side.
- nikvaes 3y agoDid you also look at the false positives, e.g., how many non-spam content was filtered by Akismet?
- dimmke 3y agoOf course. It would be bonkers not to. It just doesn't send a notification if the submission is flagged as spam and puts it in a separate folder. So I have the ability to look at every submission. I put the system into effect on August 1st. There have not been any false positives. There was even a submission to the form that was clearly a B2B sales pitch, but because it was an actual person submitting the form and not an automated system it went into the "real" entries list (I think this is reasonable. Any business is going to have to field B2B sales solicitations) I put together a few rows in a spreadsheet of legitimate submissions (with info blocked out): https://imgur.com/a/stxja1Z https://imgur.com/a/stxja1Z Here's an example of one flagged as spam by Akismet that was submitted about an hour ago: https://imgur.com/a/PmN3t80 https://imgur.com/a/PmN3t80 Overall, removing reCAPTCHA has increased the total amount of submissions to the form, but the amount of submissions actually being seen by a real person who then has to waste time reading it, identifying that it's spam and discarding it has dropped to 0.
- dontupvoteme 3y agoI wonder if it would be a decent approach to scale CAPTCHA tests based on how likely an LLM thinks a post is spam. People who WRITE LIKE THIS might just end up ACCIDENTALLY BEING FILTERED and using click as an imperative is a complete red flag. I think you could probably filter a number of these out with just regular NLP approaches/models even.
- base 3y agoAkismet is a paid service and their apis are tailored for comments. An advantage with comments is that you can just mark as spam if some contents have dubious links or keywords. An issue you have in many forms (e.g.: login form) is that there is limited data to decide if it's a real user or a bot.
- dimmke 3y agoAgreed on all points. That's why I said in my original comment: "I know these get used for more than [contact] form submissions, to stop bot sign ups etc…" I picked up Akismet because it's been around forever, and while it is paid, it is very cheap for my use case. This is a bit of an aside, but I feel like Automattic is sitting on several companies/products and not doing a whole lot with them. Akismet could be expanded into a more fully featured server side spam detection SAAS with a flexible API etc... Gravatar could be expanded into something like OpenID. Just seems like a waste to me.
- hackermatic 3y agoThe answer is something like "yes, and..." because reCAPTCHA already decides whether and how to challenge the user based on its own internal risk score.
- dimmke 3y agoBut if a server side only solution seems to work fine, why add a client side element? I ran into this because I was doing some freelance work on a website that had worked its ass off to cut loading size as small as possible. For reCaptcha to develop its risk score, you have to load it far in advance of a form - you can't lazy load it, they specifically say not to: https://developers.google.com/recaptcha/docs/loading https://developers.google.com/recaptcha/docs/loading. It also spawned its own service worker on top of adding like 300kb of page weight. The API documentation is garbage, you have to fuck around with Google Cloud to get API keys now too which is confusing. It also pollutes the global scope of the page. It's all around terrible to work with.
- ravenstine 3y agoSounds like you should write an article about this! I've never heard of Akismet so I'd be curious to see more info on your findings. CAPTCHAs today are horrible. I don't even bother with most webpages that require them at this point. In a similar vain, I don't bother with sites that send me in a never-ending Cloudflare loop just because I lock down my browser to limit what sites can and cannot do. It's particularly tyrannical when I am either sent in a loop or asked to complete a captcha for a non-interactive page.
- bawolff 3y agoSome spam is really low effort. Even a non obfuscated text image (e.g. something that can be read by teseract out of the box) still stops a surprising amount of spammers.
- chrismorgan 3y agoI added this trivial honeypot field to a site’s register-interest form in late 2021, and it has been very effective at culling spam: I had been getting an average of around one spam message a day, but after adding it it took a year and a half for one to get through, and no others have got through. <style>.pot{display:none}</style> <div class="f pot"> <label for=username><b>If you are human, leave this field blank:</b> <em>(required)</em></label> <input name=username id=username> </div> The whole field is hidden with CSS; the “if human, leave blank” instruction is thus only used for text browsers, but I still prefer to have it.
- rkta 3y agoAs a text browser user I thank you for adding that label.
- runeks 3y agoI wonder if this is filled out automatically by browsers like Chrome when doing auto fill...
- chrismorgan 3y agoI would be extremely surprised if it filled it: `display: none` makes the field not be rendered, and autofill should only fill stuff the user could fill.
- thebears5454 3y agoI think that's one of the benefits of using things like Auth0. They have thousands of companies so their heuristics are really good.