13 ms·
Fighting API bots with Cloudflare's invisible turnstile
- krackers 3y ago"Invisible" assuming you have javascript enabled and use a mainstream browser. The failure mode on these is worse than regular captchas because cloudflare won't even give you a chance to prove that you're a human, you'll just be stuck in a refresh loop.
- shortcake27 3y agoSurely people who disable javascript are used to the majority of the web being broken for them. JS is an integral part of the web, regardless of how people feel about it.
- throwitaway156 3y agoit doesnt matter if you enabled it or not, you need to allow Clownflare's guerrilla fingerprinting to be allowed access. Which is a huge downside, unless you have access to a bunch of proxy servers and the knowledge to randomize and spoof your fingerprint accurately.
- Karellen 3y ago> Surely people who disable javascript are used to the majority of the web being broken for them. Not really. I'd say >80% of sites I visit are accessible, in that I can read the content of the page I want to look at, with javascript disabled. Of the remainder, I can temporarily switch javascript back on for a tab with two clicks, if I want to read the contents enough, and I think the odds of that site having especially malicious javascript on it is slim. (e.g. mastodon instances) I can also create permanent exceptions for websites in a couple of clicks too, and there are some frequently-visited sites that I have done that for. Don't think of those who disable javascript as browsing with no javascript. Think of it as browsing with javascript disabled by default.
- ZiiS 3y agoHe implemented a fallback to a regular captcha to cover this case.
- randunel 3y agoBut that captcha doesn't work no matter how many times you solve it https://i.imgur.com/1wEy8op.mp4 https://i.imgur.com/1wEy8op.mp4
- nonameiguess 3y agoYou don't even need to do anything unorthodox. I'm using Firefox with Javascript enabled, set to block fingerprinting and cross-site cookies, and Cloudflare's bot detector regularly puts me into infinite loops. It's especially frustrating when it's the login page for a paid service doing this to me. Why on earth would I abuse your site with a bot using credentials associated with my real name and payment info?
- 1vuio0pswjnm7 3y ago"Which poses an interesting question: how do you create an API that should only be consumed asynchronously from a web page and never programmatically via a script?" Web developers use JavaScripts to make HTTP requests to API endpoints. The data is being consumed by a script, programmatically. Unless Javascripts are neither scripts nor programs. Good luck with that argument. There is a W3C TAG Ethical Web Principle that states web users can conusme and display data from the web any way they like. They do not need to use, for example, a particular software or a particular web page design. 2.12 People should be able to render web content as they want People must be able to change web pages according to their needs. For example, people should be able to install style sheets, assistive browser extensions, and blockers of unwanted content or scripts or auto-played videos. We will build features and write specifications that respect peoples' agency, and will create user agents to represent those preferences on the web user's behalf. With respect to the JavaScripts authored by web developers and inlined/sourced in web pages, web users have no control over them short of blocking them outright. As such, arguably they are not ideal for making HTTP requests to API endpoints. Unfortunately these scripts can be, and are, used to deny web users' agency.
- andersa 3y agoWho cares? He wants to provide a free service, which he is not obligated to do and costs money for him to do, for individuals to check their email. He never intended for bots to check a billion emails, and obviously doesn't want to pay for that. That should be respected, and people failing to respect it is why we see the destruction of the open web with things like remote attestation as the only way forward. Complaints like "but the open web" or "but my exotic browser" are honestly worth nothing against a potential solution for a real, pressing issue like spam requests and bots - if we want to keep the nice things we better start coming up with alternative solutions to the real problem, because corporations will decide the future for us if we don't. Ignoring it will not work out for us.
- ZiiS 3y agoIt often costs far more then money. In the this case bots were gaining data to better target future attacks. In the common spam case it costs real user's attention.
- fswd 3y agoRecently fighting with bots in a different situation. I discovered you can return code "466" in nginx, which is a special code that completely disconnects the TCP session.
- nicoboo 3y ago466 vs transparently slow the response (exponential throttling)? To avoid having auto reconnecting behavior? Maybe both.
- tlavoie 3y agoLike a Slow Loris attack, but from the server side? I like it! I've been using a mostly-Apache setup for ages, but thinking about how it might be fun to implement something lightweight for my VPS, that includes a variety of ways to mess with those sending unwanted requests. I suppose ModSecurity could get me most of the way there without having to reinvent everything.
- jeroenhd 3y agoIf you're still on iptables, you can TARPIT traffic using firewall rules that will essentially do that. nftables doesn't have tarpitting just yet, I believe. If you want to annoy SSH brute forcing bots, endlessh is a dedicated tool for SSH connections. There are other tools for other dedicated protocols as well.
- tlavoie 3y agoCool, thanks! I do use fail2ban on my VPSs fairly liberally, so filling any one log with too much noise will trigger an hours-long ban for the IP. What I liked about the application-level interference is that you can do something more subtle than a block, while still feeding them nonsense, slowly.
- fswd 3y agoMy second thought was utilizing some nodejs express reverse proxy -- with some kind of rate limiting slow down, but the attack stopped and I moved on to something else.
- nicoboo 3y agoReally interesting article and details, we are in similar cases, I would definitely consider such implementation and we'll also look at alternatives to Cloudflare. Thanks you Troy for writing and sharing your experience.
- politelemon 3y agoThe concept makes sense to me, it's reminiscent of hmac, but with the additional constraint that the "secret" key is in the open. And the server side verification is opaque so you have to hope it works. But I like that it's not captcha which has made web browsing worse over the years I'm curious to know if there have been similar open source implementations of turnstile where website operators have found ways of limiting an API call to just a browser, without captcha. Does anyone know of any?
- ZiiS 3y agoThis is as practical solution to a very real problem. HIBP is so valuable both the no API case, and the case where bots scrape the API reduces people's online security. However relying on the near universal behaviour tracking and fingerprinting of large corporations is extreemly worrying. The better Turnstile works the more like Google's proposed Web Environment Integrity it becomes. https://www.eff.org/deeplinks/2023/08/your-computer-should-say-what-you-tell-it-say-1 https://www.eff.org/deeplinks/2023/08/your-computer-should-s...
- ytch 3y agoI encountered a similar problem at work recently. The first thing come to my mind is CSRF. But it can be bypassed by parsing the token in HTML. Then I found this article[1]. Embedding a challenging with valid time in Javascript, calculating the response by obfuscated Javascript code. [1] "Bypassing YouTube video download throttling" https://news.ycombinator.com/item?id=37117338 https://news.ycombinator.com/item?id=37117338
- buro9 3y agoJust put in a multi second delay to _all_ requests to that API, humans wait but this costs bots and slows the aggregate progress. Add HTTP headers documenting the API you want them to use as a bot author will look and wonder why it's performing so badly. The difference between human speed and computer speed is so noticeable that it can be leveraged... No high bar or complex adversarial solution needed, and the side benefit is that if a bot does persist it spreads the load.
- mordae 3y agoNot sure why downvoted without much comments. I can see a possible problem in that the bots would just establish multiple parallel connections and wait.
- kevincox 3y agoThis will harm the experience for the intended user and will barely affect the bots because they will just make more requests in parallel. The user is going to be much more speed sensitive than a scraper.
- nijave 3y agoA delay would tie up TCP connections/sockets on the bots but it would add trivially compute cost. Couldn't you just increase the parallelism? I think a lot of these schemes rely on solving a CPU or memory based computation with JavaScript
- jeroenhd 3y agoThe bots send out half a million requests, who cares if they need to wait ten seconds per request. You're adding a total of ten seconds to hours of scraping.
- neonate 3y agoDid I miss something? I don't understand what the "non-interactive challenge" is here: > The widget takes responsibility for running the non-interactive challenge and returning a token I get that a widget can return a token and then you trust the token and do the rest. But what determines whether a caller receives a tokens or not? That is, what's the actual challenge? That is, how is the distinction between bot vs. real human user actually getting decided?
- allisdust 3y agoProbably a form of proof of work which is costly for the bots but fair on normal users (combined with normal fingerprinting and IP check which determines how hard the challenge should be for a request)
- RobDukarski 3y agoIf you're just pinging the endpoint without providing the necessary header then you are effectively a bot in this case. If you go through the motions to fill out the form and then submit it you're at least not explicitly telling the service you are going against its wishes by pinging the API directly. As for the challenge, that seems to be within Cloudflare's implementation which then returns the token to submit with the form. HIBP then verifies the token to make sure its a match before checking for the data and sending a response.
- wewxjfq 3y ago> the unsolved challenges are when the Turnstile widget is loaded but not solved (hopefully due to it being a bot rather than a false positive) So people employ these measure and have no clue whom they filter. Reminds me of the online shops who block you because you click on the products too fast. Congratulations, you lost a customer to keep the CPU utilization at 20%.
- Rebelgecko 3y agoI'm trying to buy a car, which involves going to a bunch of dealer websites and seeing what they have in stock. Usually after checking 2-3 dealers I'll get a Cloudflare error and can't access any car dealer websites for a day or two (or I can keep going if I use a different browser, or just Incognito mode). I guess "Turnstile" might be what I'm running into?
- mtsr 3y agoI run pihole, my own DNS resolver and some other stuff and I get the cloud flare checking your browser page on every single site using cloud flare. Fortunately I haven’t seen the longer blocking you mention.
- leesalminen 3y agoI’ve been running pihole with my own resolver for >5 years and don’t have this issue. I run into it much more frequently when behind CGNAT.
- cinntaile 3y agoThis was posted last week. It still reads like an ad and it very likely is one.
- jeroenhd 3y agoHIBP receives a lot of commercial support from Microsoft and Cloudflare, and in turn provides their API to organizations like government agencies and Mozilla. Half of the articles about it are "this company gave me this free/cheap thing, here's how I implemented it" and I can't feel too bad about it. What's weird is that the current implementation is broken. Once you do get a CAPTCHA to fill out, the redirect will fail and you end up starting over, only to get a new CAPTCHA page. That's rather unfortunate.
- lionkor 3y agoThat peak is around 400 req/s, right? I would expect the usual solution to be to just rate limit each user to some reasonable limit, and then 429 if theyre over. I feel like 400 req/s should be absorbed, especially since my 5€/month VPS can handle about 2x that, sustained. I might be missing something here, but that just doesnt seem like a peak big enough to warrant degrading user experience. Sadly I'm often categorized by websites as "probably a bot haha get fucked", and its lost sites hundreds or thousands of $$$ worth of revenue over the years, just from me alone.
- nilsherzig 3y agoThe backend has to search through TBs of data per request
- 3y ago
- hoseja 3y agoNot a fan of nasty fingerprinting tricks honestly.
- kurtoid 3y agoIt's feels more like they're cavity-searching your browser, not just fingerprinting
- SCUSKU 3y agoI am most likely misunderstanding this, but why can't you just use browser automation to then generate the turnstile token? For example, just use Playwright/Selenium/Phantom to generate the token and then use that in an API call? Otherwise, great write-up! And excellent service!
- ahofmann 3y agoThis will cost you time and most likely money. And that's stopping a lot of attackers. And if you don't stop them, you at least slowed them down big time. This could also be enough to make the attack useless.
- didntcheck 3y agoAnd in addition to the implicit proof-of-resources of forcing attackers to run a bunch of Chrome slaves, there' are also explicit POW/S challenges in the code, according to the article. It's quite an old idea [1], to add a cost which is trivial for users but a significant overhead for spammers [1] https://en.wikipedia.org/wiki/Hashcash https://en.wikipedia.org/wiki/Hashcash
- michaelt 3y agoCloudflare's code attempts to detect browser automation is happening. For example, desktop computer but clicking things without moving your mouse? Suspicious. Say you're a phone, but have desktop computer fonts installed? Suspicious. And suchlike, the precise methods are the results of a cat-and-mouse game. If these heuristics identify your browser as suspicious, they either show you an interactive captcha, or they just refuse your request.
- nijave 3y agoYup, some of these are pretty invasive like opening websocket connections to localhost to try to find daemons running on the clients machine (I think eBay and maybe others were doing this)
- yjftsjthsd-h 3y ago
- bArray 3y ago> That's a 91% hit rate of solved challenges which is great. That remaining 9% is either humans with a false positive or... bots getting rejected If I meet a "human check", I quickly decide whether it is worth me solving it, or just close the tab. I could imagine 9% of people just giving up. Some of these CAPTCHAs require you to find 20 fire hydrants on 3 different rounds of tiles, just to fail you anyway. We have loads of data on websites keeping user's attention [1], this also seems to apply to CAPTCHAs. Besides, I think it is now well known that AI is fully capable of solving CAPTCHAs. [1] https://www.nngroup.com/articles/response-times-3-important-limits/ https://www.nngroup.com/articles/response-times-3-important-...
- jeroenhd 3y ago> Besides, I think it is now well known that AI is fully capable of solving CAPTCHAs. That's the biggest downside of modern AI, and I fear the web will only get worse because of it. If we can't figure out how to patch CAPTCHAs against bots, remote attestation will become the norm.
- bArray 3y agoI have also thought about this extensively, but haven't really come to any real useful insights. A few ideas I have considered: 1. Embrace the bots and get each request (with response) super lightweight. Anything you can pre-compute, pre-compress, pre-cache is great. I've used this successfully for a small service that can scale significantly. 2. Make the cost of interacting with your service computationally expensive. For example, you could send off a problem to be solved which becomes itself a token to make one interaction. There are several problems that are computationally expensive to compute, but easy to verify. 3. Make the cost of interacting with your service require sending a significant payload - the idea being that if they launch many requests from a single network, they saturate their network. If to watch a 100MB Youtube video you had to send 1MB of random data via UDP (used to fingerprint), I suspect people abusing your service would soon find they experience dropped packets. If they struggle to send 1MB of random data, there's a good chance they would have trouble downloading 100MB of data. 4. A lot of these AIs falsify information to appear plausible. You could abuse this to ask questions, some real some false, and brief the user to answer randomly on the nonsensical questions. For example, "How many connections are there in a tripoduplex?" Something like chatGPT may see tokens for "tri" and "du" and output 3 or 2. There would also be a way to do this with images, i.e. "Select all of the cats in the image and press done", where they are all some weird trip of images. These are just some ideas and there are obvious flaws in some of them.
- mrieck 3y agoCloudflare doesn’t like my VPN so I get a lot of their challenges. Now whenever I see one I just close the tab. Let’s say you run a SaaS using Cloudflare. You may be extremely happy you block 10000 bot requests for every false positive that’s a real human. But let’s say that false positive was a potential customer that would only have paid if they weren’t blocked, and now you just saved less than a penny in server costs to lose hundreds of dollars of lifetime value from that customer. Sure if you run a free service use Cloudflare. Give in to centralizing the web more, supporting more censorship, and annoying the hell out of a ton of people in the process. But if you’re making money, I don’t see why you wouldn’t have authentication tied to paying users for anything of value, or think of bot traffic as a cost of doing business.
- kredd 3y agoCost of business. Depended on your SaaS, you could be saving more money blocking the requests to those bots than acquiring one customer. It’s a trade off you have to deal with when you get to a certain scale.
- jeroenhd 3y agoCloudflare isnthe solution to VPN companies allowing bot accounts to taint their IP address. It sucks as a real person using a VPN, but having your website be overwhelmed by bots suck more than a few VPN users having trouble using your website. If Cloudflare didn't exist, websites would probably just block VPN IPs like streaming services do.
- nijave 3y agoDepends on the application. I worked on one where new customers cost upgrades of $8 related to cost of reporting requirements (similar to background checks). Letting 10k bot requests supply stolen identities was incredibly costly. Similarly, say an e-commerce business releases a limited edition product. Many users won't end up getting it anyway so blocking a few users is usually a much better experience than letting bots buy the product for resale later. On the other hand, it's absolutely infuriating when blogs/search engines come up with these.
- IMTDb 3y ago
- hkt 3y agoMy isp uses cgnat, so I see these all the time. The two tier internet is here and I hate it.
- jeroenhd 3y agoI think it's been here for much longer. Back in the day, IP addresses would just get blocked; fail2ban was a recommended tool for any web server back in the day. Getting flagged as a bot sucks (I've experienced it for a few days myself) but the modern CAPTCHA solutions are a lot better than the "server did not respond" days from before. ISPs still failing to implement modern networks and sticking with broken workarounds like CGNAT are as much to blame as the bots tainting their CGNAT IP addresses.
- hkt 3y agoNo doubt, but at least in the UK nobody is forcing ISPs to modernise their infrastructure because competition seems to be entirely on price rather than service.
- superkuh 3y agoI dont' want to victim blame here, but if your "ISP" uses cgnat it's not an ISP. It's a web service provider. You should probably get a real ISP. That said, I have a real ISP, comcast, but comcast does MITM attacks on it's users so I have to tunnel everything through various VPS I rent. Which of course means I get hit with the same cloudflare blocks. And the invisible javascript ones just go in loops no matter how many times I complete the visible side. I just close cloudflare hidden sites' tabs' now. The problem is more and more of the web is hidden behind their computational paywall... even academic journals now.
- yjftsjthsd-h 3y ago> I dont' want to victim blame here, but if your "ISP" uses cgnat it's not an ISP. It's a web service provider. You should probably get a real ISP. Okay, so that... that is victim blaming. If you don't want to do that, stop doing it. Besides, do you really think someone in this situation can just get a real ISP? I mean, maybe they're in a competitive market and just managed to pick a bad option, but it's unlikely.
- philo23 3y agoInteresting that both the client side API and server side API for Cloudflare's turnstile seem to match Google's reCAPTCHA nearly exactly, which works in pretty much the same way with the exception that you can't configure it to _never_ show a visual captcha (in rare cases the v2 Invisible reCAPTCHA will still show the "select all the X from the images below" dialog) Even down to the API endpoint and JS API names. https://www.google.com/recaptcha/api/siteverify https://challenges.cloudflare.com/turnstile/v0/siteverify grecaptcha.render({ callback: function (token) { ... } }); turnstile.render({ callback: function (token) { ... } }); As soon as I saw the examples I recognised the names, I guess it's designed to be a drop in replacement? https://developers.google.com/recaptcha/docs/verify https://developers.google.com/recaptcha/docs/verify https://developers.google.com/recaptcha/docs/invisible https://developers.google.com/recaptcha/docs/invisible
- datguyfromAT 3y agoi think it was intended as a replacement; made it easier for me to give my clients the possibility to choose between different captcha services while i only have to code one (with some minor quirks) implementation
- randunel 3y agoHow HIBP works for users navigating from third world countries: https://imgur.com/a/K5z1X2R https://imgur.com/a/K5z1X2R It's a long .gif file which shows that cloudflare's website loads just fine, but HIBP is unusable. Thanks Troy and Cloudflare for making this (free) service unusable. It's free, so I shouldn't expect that it works, anyway. Chrome on Linux and no VPN, fwiw.
- mananaysiempre 3y ago>> https://imgur.com/a/K5z1X2R https://imgur.com/a/K5z1X2R > This post may contain erotic or adult imagery. Yeah it’s fucked alright.
- randunel 3y agoThe "Mature (?)" checkbox is unticked, and there doesn't seem to be a different setting for it.
- jiofj 3y ago[flagged]
- randunel 3y agoThere's a lot of English and American bias on the internet, I don't know if "stop being poor" is the solution, but I'm definitely trying.
- jiofj 3y agoThis has nothing to do with "English and American bias", this is simply the result of most traffic coming from your country being bots because of the amount of compromised machines that there are there.
- 46Bit 3y agoThat looks like a misconfiguration by HIBP
- 3y ago
- kimburgess 3y agoFor this style or abuse mitigation I’m always surprised that HashCash [1] or similar simple, locally implemented proof of work mechanisms aren’t more common. This can be implemented in a way that remains transparent (albeit via JS), poses little impact on ‘good’ users, but protects against a lot of traffic patterns that may be undesirable. The cost can be scaled to match infra capability and the challenge can be a combo of the request data and time. Valid windows for that time can then be synced with cache validity which removes the need to keep tabs on any state. For those deeper in this space. What am I missing here that prevents this from being the norm? [1]: http://www.hashcash.org/ http://www.hashcash.org/
- michaelt 3y agoIt turns out some of the abusers are using 'botnets' of thousands of virus-infected home PCs. So they've got thousands of CPU cores available for proof-of-work challenges, legitimate residential IP addresses, and so on. Meanwhile, plenty of the legitimate users are using 5 year old budget android devices, so you'd better not make that challenge too hard.
- nijave 3y agoYeah, there's lots of these floating around sometimes called "scraper service" or "residential proxy". Not sure if it's still around, but one of them enlisted machines by paying users to install a browser extension.
- jeroenhd 3y agoThere was one famous free VPN service that worked like this. You install the addon, get a free VPN for a certain amount of traffic, and while your browser is open other people will be able to browse from your IP (and access your home network, of course!) Making the browser deal with PoW challenges is only a small price to pay for what is practically a free VPN. It works great, until your entire home IP starts getting CAPTCHAs all the time, and because users don't know any better, they start blaming that darn Google/Cloudflare/Microsoft for claiming they're a bot.
- mike_hearn 3y agoI did a lot of work on this many years ago at Google. As the article says, it can work well and be minimally invasive for users (they need to run JS but that's a much lower bar than solving complicated CAPTCHAs). There are several services like Turnstile. I'm an advisor to Ocule [1] which is a similar thing, except it's a standalone service you can use regardless of your serving setup rather than being integrated into a monolithic CDN. It's a smaller company too so you can get the red carpet treatment from them, and they aren't so aggressive about blocking VPNs and privacy modes because their anti-bot JS challenges are strong enough to not need it. They're ex-reverse engineers so know a lot about what works and what doesn't. Their tech may be worth looking at if you're concerned about over-blocking. The mention of Turnstile using proof of work/space is a bit puzzling/disappointing. That stuff doesn't work so well. There are much better ways to create invisible JS challenges. The core idea is to verify you're in the intended execution environment, and then obfuscate and randomize so effectively that the adversaries give up trying to emulate your code and just run it, which can (a) lead to detection and (b) is very slow and resource intensive for them even if they aren't detected. Proof of work/space doesn't prove much about the real execution environment. BTW, the author asks what proof of space is. It's where you allocate a huge amount of RAM and then ensure the allocation was actually real by filling it with stuff and doing computations on it. The goal is to try and restrict per-machine parallelism by causing OOM if many threads are running in parallel, something end users won't do. Obviously it's also a brute force technique that can break cheaper devices. [1] https://ocule.io/ https://ocule.io/
- arsome 3y agoI think the idea is they ramp up the difficulty of the proof of work when a user is suspect by their other tests, after ramping it up to bot levels it causes those requests to become more costly even if they're not blocked.
- mike_hearn 3y agoIt hardly works. PoWs can be hyper-optimized in bot code compared to in browser JS, and your PoW has to be tolerable even for slow old devices so it can't be too demanding anyway. It also assumes a very smooth gradient of suspiciousness. When using JS challenges though, you can be often working with binary signals (at least, that's what I was able to get). So then there's not much "maybe so/maybe no" about it. Either you detect a bot with 100% confidence, in which case you just drop the banhammer. Or you don't spot it and have to let it through.
- johnklos 3y agoThis reads like an ad. There are much better ways to deal with bots than to let Cloudflare further bifurcate the Internet.
- varun_chopra 3y ago> There are much better ways to deal with bots Such as?
- johnklos 3y agoFor something like this service, simple rate limiting per IP / netblock, along with TCP/IP fingerprinting for VPN endpoints and such, could be very effective. I run histograms for connections per netblock on my email servers, and even removing only the most egregious attempts at abuse almost empties my logs. On the other hand, Cloudflare has issues with tons of less popular networks, with VPNs, with less affluent countries, with non-mainstream OSes and browsers, et cetera, all of which ends up punishing many people in ways that are completely disproportionate to the amount of abuse avoided. It reminds me of the quote from fortune(6): As far as we know, our computer has never had an undetected error. -- Weisert You don't know how many people Cloudflare has marginalized because you don't see their visits.
- thiht 3y ago> There are much better ways to deal with bots I'm interested, do you have some resources to share?
- johnklos 3y agoSee https://news.ycombinator.com/item?id=37425473 https://news.ycombinator.com/item?id=37425473 It's just referencing an example, but creating usable tools really isn't hard.
- 1vuio0pswjnm7 3y ago"And then, unsurprisingly in retrospect, it started to be abused so I had to put a rate limit on it. Problem is, that was a very rudimentary IP-based rate limit and it could be circumvented by someone with enough IPs, so fast forward a bit further and I put auth on the API which required a nominal payment to access it." If one does not know how to limit based on number of HTTP requests, then is one really qualified to set up a "public API". It is well-known that IP-based limits, i.e., blocklists, do not work. (While allowlists are common, e.g., academic journals, Verisign zone files, etc.) We cannot blame the public because someone does not know how to configure a proxy to limit number of HTTP requests per IP, e.g., in a 24-hour period. Here, the public is penalised by asking for credit card numbers because the site operator does not know how to count HTTP requests. One would think people who do not want their details in a data breach that random people can download probably would not want to give their details to some random person whose website becomes popular. But it seems they do. HIBP never made any sense. "Send me your private info and I will check and make sure it has not been leaked. Too late. You just leaked it to me." This sort of obvious blunder had to be fixed. Still too late for anyone who used it before HIBP was "fixed". Why place any confidence in someone who cannot spot these issues. Here, he struggles to implement an API. Stupid websites can become popular. It happens. Many websites have sought to exploit data breaches. What better way to collect working email addresses that people care about than to let people submit them to you to check against a dump of a data breach. Unless they download the dump themselves, they have no way to confirm you actually checked anything. And if they did download the dump(s), then there is no reason to submit anything to you via an "HIBP" website. If people are getting charged for API access, and submitted personal info to HIBP, then they should get some enforceable terms in return. If you collect peoples' information then you are liable for the damage that may result if HIBP is breached. Doubtful the API customers get any such protections.