4 ms·
Burn the witch! Lets read through that page for a second though: Drop support for obsolete HTTP versions Doesn't seem like that's going to cause much issue
by daenney 4y ago
Burn the witch!
Lets read through that page for a second though:
Drop support for obsolete HTTP versions
Doesn't seem like that's going to cause much issue for any legitimate client from the past 10-20 years. He only recommends blocking HTTP 0.9/1.0, which fair enough
Append a #hash to the form’s action URL
Hah. Clever man. I don't see how this is going to stop any legitimate user from loading your website or submitting the form, but I can see how it might frustrate bots.
Include a hidden prefilled form field
This is just standard practice to mitigate CSRF.
Verify the Host and Origin request headers
Yes. You should be doing that.
Set a test cookie and verify it gets included in the submission
Another CSRF trick.
Swap the name attributes in the name and email fields
This one's a little user hostile to folks who use assistive devices like screen readers. But still won't prevent you from accessing the site in the first place.
Verify the POST/Redirect/GET (PRG) chain
As noted by the author, might cause some issues but again, won't stop anyone from loading your website.
Block ancient versions of common browsers
Alright please just don't do this. UA blocking is gross and might prevent access through specialist software. But he also calls this out himself.
I strongly discourage you from blocking or discriminating against unknown or uncommon browser User-Agent request headers
All in all, with the exception of UA blocking I don't see how any of these mitigations would result in users not being able to access said website, or having their loading times drastically increased.
- nijave 4y agoAll that stuff is easily defeated by automated browsers anyway (i.e. selenium)
- mh- 4y agoYes, but those automated browsers are much more expensive to operate than simple HTTP clients pretending to be browsers. It's an arms race/defense-in-depth situation. If someone truly wants to automate your site in a targeted fashion, and it's profitable for them to do so, you'll have to invest a lot more in stopping it (and decide how much of it is worth stopping).
- Aperocky 4y agoEven youtube fails with yt-dlp going as far as a internal python file that parses javascript and execute them.
- zxcvbn4038 4y agoThere is nothing they can do, anything they can imagine to use as a signal for a bot user can be made to look legitimate. There is no way to win. My college English professor was so obsessed with beating cliff notes that all of tests were hyper focused on the most obscure details he could dream up. It was just impossible to maintain a full course load and memorize what was the 1st, 3rd, 5th, 7th , last, etc word on every page and which character spoke it. Did the sentence contain any commas? How many times was the nurse mentioned in chapter X, etc. Google can always follow his lead and make their data impossible to access, but impossible doesn’t increase ad views, so they will never do it. People using ytdlp is just the cost of doing business.
- LinuxBender 4y agoHe's quite tame compared to me I suppose. I block anything that is not HTTP/2.0 which currently knocks out all the bots and all crawlers except Bing. But I just have hobby sites these days. Nobody would notice or care if my sites went offline. Using NGinx as an example: if ($server_protocol != HTTP/2.0) { return 403 'Nope'; } Another thing I have found useful to drops some bots is to become invisible to them. Many of the poorly written scanning tools do not properly set MSS for reasons I still don't understand. I use this to my advantage. Using IPTables as an example: /sbin/iptables -t raw -I PREROUTING -i eth0 -p tcp -m tcp --tcp-flags FIN,SYN,RST,ACK SYN -m tcpmss ! --mss 420:16384 -j DROP Any TCP packets setting a very low or high MSS or missing MSS will be silently dropped. I drop about 35K packets per host per day on average. This also drops hping3 floods.
- judge2020 4y agoIMO blocking bots isn't too big of a concern, the problem is when a dedicated attacker realizes you serve valuable data (in your HTML). Next thing you know, they're running puppeteer or a similar remote controlled browser to scrape your site, which is both undesirable in itself and the scraper might overload your site/database by scraping with no internal parallel request limit. If you're not a startup with an unlimited early cloud budget, it can be costly if you want to handle both bot usage (including official API-based or scraping bots) and regular users.
- toast0 4y ago> Many of the poorly written scanning tools do not properly set MSS for reasons I still don't understand. MSS issues attract me like a moth to flame [1], so let me ask some questions. It looks like this is dropping syns with MSS over 16384??? That is indeed a pretty crazy high number. 9000ish seems reasonable for someone on a jumbo network without a mss clamping router, but above that is someone weird for sure. Under 420 seems unlikely too, but technically acceptable, but sure, I'd drop it. In theory, a proper OS will send several SYNs with MSS, then assume your server doesn't support TCP options and send you a SYN with no options. Going to take a while, but if someone legitimately has a mss less than 536, their internet is probably pretty junky anyway, so ok, seems fine. [1] I just built a browser based pmtud test site, http://pmtud.enslaves.us/ http://pmtud.enslaves.us/
- ceejayoz 4y ago> This one's a little user hostile to folks who use assistive devices like screen readers. As long as you're using a <label> or aria-label attribute, that shouldn't be an issue.
- d2wa 4y ago(Author here.) I am. There’s plenty of accessibility labels in place. It’s literally just the name attributes. No user ever sees this, whether they’re using accessive technologies or not. It only confused bots that assumes that the field named email is for the email address.
- d2wa 4y ago>> Verify the Host and Origin request headers > > Yes. You should be doing that. (Author here.) If I remember correctly, his browser of choice predates the Origin header.
- daenney 4y agoAlright well fair enough. Looks like that’s only been supported since Fx 70 released somewhere in 2019. So maybe don’t do that depending on what you intend to block. But then again it’s been 3 years also. In general though the whole tone of parent of “I am owed access to someone else’s computer system on my and my terms alone” just doesn’t jive with me. It’s also not remotely comparable to Cloudflare’s approach of sitting in the middle snd then appropriating end-user compute resources without their consent to fuel their business.