6 ms·
Who cares? He wants to provide a free service, which he is not obligated to do and costs money for him to do, for individuals to check their email. He never in
by andersa 3y ago
Who cares?
He wants to provide a free service, which he is not obligated to do and costs money for him to do, for individuals to check their email. He never intended for bots to check a billion emails, and obviously doesn't want to pay for that.
That should be respected, and people failing to respect it is why we see the destruction of the open web with things like remote attestation as the only way forward.
Complaints like "but the open web" or "but my exotic browser" are honestly worth nothing against a potential solution for a real, pressing issue like spam requests and bots - if we want to keep the nice things we better start coming up with alternative solutions to the real problem, because corporations will decide the future for us if we don't. Ignoring it will not work out for us.
- ZiiS 3y agoIt often costs far more then money. In the this case bots were gaining data to better target future attacks. In the common spam case it costs real user's attention.
- andersa 3y agoThere definitely is an alternative solution here that preserves people's freedom to use whatever browser they want, but I'm not sure if anyone would like it. Web services are intended for humans to use them, and all abuse also ultimately comes from humans directing computers to be abusive. Thus, rather than attaching a computer's temporary identity (an IP address) to the request, we should be attaching a human identity to it. Note here, that I don't care whether the request comes from John Smith in New York. I care about being able to ban you from the service if you are abusive, no matter how many computers you have at your disposal now or in the future. There's lots of downsides I haven't got solutions for, such as, how do we stop sites cross referencing to reliably discover all users real identities (Google analytics would love this!), and so much more. But we'll have to give up something - if we do nothing we give up everything, maybe giving up only absolute anonymity could be preferred.
- CameronNemo 3y agoIP reputation doesn't seem to be much of a thing right now, but it might be the only way forward for a spyware-free web. You can have a group of people sharing a block and collectively responding to abuse reports to slightly improve privacy. That at least shields you somewhat from the big tech firms. Essentially a large number of virtual private micro-ISPs.
- kmlx 3y ago> IP reputation doesn't seem to be much of a thing right now disclaimer: acquaintance works in the spam business. someone i tried to steer clear of, but was fascinated of what they told me. IP reputation is a huge thing in the spam world. they pay top dollar for residential US/UK/etc IPs which they can then use for spamming others. and we're not talking one or two IPs. we're talking 100s of thousands of IPs being transacted daily globally. all for spam. for anyone interested in seeing how low we've come, they have a huge convention in Las Vegas. one should visit it to better understand a field that has been growing like crazy but so few know it. apparently everyone is preparing for a big boom next year, with 10s of $B ready to be deployed.
- andersa 3y agoIP reputation is absolutely a thing, as I understand it it's a large factor in Cloudflare deciding to put you in an infinite loop (or trigger the protection at all).
- didntcheck 3y agoYep. Even within a single VPN service it's often noticeable how many more CAPTCHAs (or outright blocks) you get via different servers
- CameronNemo 3y agoOK, but remote attestation isn't really a thing on the web (right now). And you can script a browser UI, anyway. The point is to not get hung up on the client as a security boundary (it isn't, can't be, won't be), but to focus on the actual harm -- excessive use of the limited resources provided. And you have to frame your security posture as rooted in the server side throttling, heuristics, et cetera. Flipping out because the client isn't what you expected isn't going to help (determined attackers can look like "vanilla" clients), and is just going to harm the long tail of actual users who are not bots.
- true_religion 3y agoIt will also harm the short tail of bots who aren’t run by dedicated attackers.
- jeroenhd 3y ago> OK, but remote attestation isn't really a thing on the web (right now) I mean, it's built into Safari and Cloudflare supposedly uses it for its CAPTCHAs: https://developer.apple.com/news/?id=huqjyh7k https://developer.apple.com/news/?id=huqjyh7k It's not as invasive as Google's attempt to circumvent ad blockers, but it's still a remote attestation system implemented in the wild already.
- Obscurity4340 3y ago> built into Safari Is it really built into Safari, or are you referring to the iCloud option that "privately" helps reduce Cloudflare demands?
- jeroenhd 3y agoPrivate Access Tokens are built into Safari. Apple gives their devices a certain amount of tokens, and Cloudflare validates them. If Apple doesn't like your device, you won't get any more tokens. Cloudflare also hands out tokens if you install their browser addon. The two companies are working together to make this an official web standard, but I haven't heard about it for a while. Maybe they're just laying low after seeing the blowback on Google's (worse) attempts at attesting devices.
- davidspiess 3y agoI agree. Our public APIs are also massively queried. The number of queries is out of proportion to the legitimate traffic. Rate-Limiting does not work, because of the volume of different ips they send against you in parallel. Our servers are not designed for such peaks. What other choice do we have but to block them.
- afandian 3y agoDid you consider API keys? Via email? That’s at least a natural throttle.
- davidspiess 3y agoIt's a public price API. There is no authentication required to list prices.
- zb3 3y agoCould you explain why that API is meant to be public and what benefits it brings to human users? What kind of data does it return?
- batch12 3y agoIn addition to blocking/throttling, I have my services provide bad data to the abusing clients.
- didntcheck 3y agoAs in "valid" but false data? Please don't. If you really don't want to indicate rate limiting explicitly, then perhaps return an invalid body, or reset the connection or similar. False positives detecting humans as bots are very common, and even rate limits are often set well within human interaction limits. E.g. more than once I've triggered 429s by opening several e-commerce product pages in new tabs for me to ctrl+tab through and filter down. I also tripped a LinkedIn anti-automation system since I was looking through quite a lot of profiles on my first day to add people - luckily they handled this well, with a clear message explaining what was going on and support reaching out to me proactively (and lifting the restriction after a few hours)
- mschuster91 3y ago> if we want to keep the nice things we better start coming up with alternative solutions to the real problem Indeed. To state the obvious: the real problem are bad actors - no matter if they are nation states, cybercriminals or people running compromised devices - and their accomplices such as ISPs not responding to abuse reports. As long as we don't get that under control (say, by threatening to cut offender countries and ISPs from the Internet and SS7 phone networks) we'll have to continue whack-a-mole'ing.
- outofpaper 3y agoI still don't get why not just implement some basic usage rate limiting. Go two fold limit the rate to 2x or 3x the fastest they can manually use the service and limit users to only being able to burst like that for a bit. You're going to have zero horid scraping once that's in place.
- andersa 3y agoBecause it doesn't work. Malicious users can circumvent the rate limiting by using botnets. If requests were somehow tied to the identity of the human operator rather than the particular computer they used, then yes, rate limiting would be all we need.
- Kiro 3y agoThat doesn't solve the problem with bots in games. They are not hammering the API, they are automating things that shouldn't be automated and drives away non-bot players.
- oefrha 3y agoDedicated scrapers have been using IP pools for more than a decade. Search for (residential) proxy pools and you’ll find a million vendors selling access. IP-based rate limiting stops cheap skids and screws over people behind CGNAT, that’s all.
- zb3 3y ago> He wants to provide a free service, which he is not obligated to do and costs money for him to do, for individuals to check their email. Not a good choice, why not make this a hashed db that could be distributed freely? Why are bots meant to be excluded? This is bad UX, I'd want to check my emails periodically in the background, but this "anti-bot" measure is meant to make me unable to do so and then demand payment. So this is openly a fight. > are honestly worth nothing against a potential solution for a real, pressing issue like spam requests and bots That is not an issue for me at all. Ok, let's face it - maybe I'm just on the other side. For me scrapers are very useful because they reduce the price needed to access the data as they introduce competition. For example price comparison sites are very useful.
- jeroenhd 3y agoSurely smart bots would call the API (https://haveibeenpwned.com/API/v2 https://haveibeenpwned.com/API/v2) rather than try to scrape the web form. HTTP calls to a web endpoint are much easier to scrape than whatever the website frontend decides to call. All you need to do is parse the Retry-After headed for the 429 error code and you can pretty much query away without worrying about CAPTCHAs.
- RobDukarski 3y agoIf the service even provides such header or 429 status code at that. They could provide the 418 status code for fun. In this case – without checking the form implementation – I can assume that it provides a JSON response rather than a text response meant to be thrown into the DOM. Grabbing data from JSON is naturally easier but you could use a DOMParser for text content itself if it's sufficiently consistent. The other thing about "waiting" is that bots may not want to do that, or maybe some sort of deadline is sooner than such waiting would allow. To me, requiring a unique key to be input with the search, that is created after X amount of time (both provided by the initial response and increasing exponentially for subsequent requests) seems like it could be sufficient. If the next request is sooner than X then block them for Y amount of time for attempting to bypass allowable behavior. Allow like 5-10 emails then implement the wait-based functionality so that most non-bots would be fine. After all we're talking about blocking an endpoint only ever meant to be used via the website by actual users, typically they are not trying to check thousands of emails super fast.