55 ms·
Avoiding bot detection: How to scrape the web without getting blocked?
- abadger9 5y agoI'm a lead engineer on the search team of a publicly traded company who's bread and butter is this domain. I was curious about this list, it candidly misses the mark- the tech mentioned in this blog is what you might get if you hired a competent consultant to build out a service without having domain knowledge. In my experience, what's being used on the bleeding edge is two steps ahead of this.
- ryandrake 5y agoDo you have any factual corrections? Your post reminds me of those "I'm getting a kick out of these replies" copypasta--declaring someone wrong and claiming authoritative knowledge, but without actually correcting any of the errors of fact.
- ComputerGuru 5y agoYes, but it’s also understandable why they wouldn’t want to expand.
- abadger9 5y agoSorry about that, we're unable to discuss our projects with adjacent teams within the company. What you're saying is a valid frustration, the motivation behind the original comment was to put a thermometer on the repo.
- gwbas1c 5y agoWell, don't go getting our hopes up!
- fragmede 5y agoAdjacent teams? It makes sense not to discuss it with totally unrelated teams, but a large Internet company wants to have its cake and eat it too - scrape all of the worlds websites, while simultaneously denying others from scraping its results. Limiting communication between adjacent teams (ie the scraper team and the bot blocker team) really seems like it would be a hinder rather than help.
- jsnell 5y agoAbuse is the kind of problem area where anyone seriously working on either side will be unlikely to go into the details. It's the same whether it's blocking scraping, stopping spam, preventing account takeovers, or detecting payment fraud. In this specific case, the people wanting to detect bots want to avoid having their signals burned, the scrapers don't want the defenders to know which signals they're able to cloak since it will spur new signals development. So what gets disclosed publicly is just the really simple stuff. It's kind of sad. There's a huge pipeline for getting people up to speed on security engineering, since there's a lot of incentive for everyone in the ecosystem to share information as publicly as possible except for relatively short responsible disclosure windows. In contrast, the only way to learn abuse engineering is to happen to work in an organization with an abuse problem (and live with the frustration of an abnormally long ramp up period), or to go black hat. And likewise it's quite hard for the good guys to actually learn from each other, since they're spread across so many companies and it's thus hard for them to exchange information on what works and what doesn't.
- ChuckMcM 5y agoNot the OP but when I was running Operations at Blekko (a search engine) I spent part of my time dealing with scrapers. When treated like a puzzle it can be really interesting. So I thought I'd share a few tidbits. 1) We did a simple 'speed' test, how many queries per second were coming from an IP, and auto-ban on the limit being exceeded, started at 100qps and watched as the traffic moved down to 99.5qps. Pushed to 10qps and watched the traffic follow it down. Even at 3qps you would get traffic at bit over 2 qps trying to limbo in under the limit. 2) At that time, lots of people who highjacked browsers with toolbars sold scraping as a service to third parties. Their toolbar would check in to see if it should do a query and it would launch a query and return the results without the user even knowing. One company, 80 legs, was pretty up front about their "service", SEO types would use it to scrape Google results to see how their SEO campaigns were doing. 3) The majority of the traffic had criminal intent, looking for metadata on web pages to indicate they were running an unpatched version of some store software or had sql injection bugs. These would often come from PCs that had been compromised for other purposes or "zombie" PCs. We could rapidly map out these networks when we got 100 queries from 100 different IPs looking for "joomla version x.y",p=1 through "joomla version x.y",p=100. We briefly played around with sending them official looking SERPs but all the links went to fbi.gov though an obfuscator. One of more effective strategies was to field a "black hole" server, basically it was an http server that answered like you had gotten hold of it but then it never sent any data. With some simple kernel mods these TCP connections were silently removed on our end so they took no resources and the client would wait basically forever. We ack'd all keep alive packets with "Yup, we're here." so they just kept waiting and waiting. It really was a never ending game. We mass banned an entire Ukranian ISP because out of billions of queries not a single one was legitimate.
- Scoundreller 5y ago> One of more effective strategies was to field a "black hole" server, basically it was an http server that answered like you had gotten hold of it but then it never sent any data. With some simple kernel mods these TCP connections were silently removed on our end so they took no resources and the client would wait basically forever. We ack'd all keep alive packets with "Yup, we're here." so they just kept waiting and waiting. Mailinator would do a similar thing with their custom email server hardware. Since they didn't really use sockets in the traditional sense, they were happy to give slooow replies and never disconnect "bad" connections.
- gzer0 5y agoI have a considerable amount of experience in the industry. Some of these so-called "advanced" techniques: * We use our own mobile emulation software (similiar to bluestacks). Turns out, mobile helps with a lot of things (below). * We use mobile IPs only. Mobile LTE data users are behind CGNATfor IPV4. You can't block one ip without possibly blocking hundreds of innocent IPs using the same exit point. * All you need is a new useragent and browser fingerprint; combined with emulation + mobile IPs, there's really no easy way for companies to block this. * With the advent and ease of virtualization; we avoid using any headless browsers. Seriously, if you can, never use headless. This should be close to rule number one for anyone looking to operate any kind of scrapers. All of our scrapers are run in isolated virtual instances with full mobile browsers. * We can easily reset our device identifier, device carrier, simulated SIM information, and especially important is the Google advertising ID that is set per device; the list goes on. The key here is #1, our mobile emulation software. * Our automation scripts are a combination of human recorded set of actions which we then perfected and can run in certain loops (for some of our data).
- AlexanderTheGr8 5y agoThis list is pretty interesting. If you don't mind me asking, what do you work on that you requires such sophisticated stuff? Also, does this work only for browsers or also for mobile apps? I have always assumed that it is always theoretically possible to get data from browsers (very extreme resort is save the browser page / (screenshot + computer vision)); but it can be impossible to get data from apps (especially ios). Are my assumptions correct? Also can you explain mobile IPs more? If they are such a big vulnerability, why is there no potential solution to them?
- kingcharles 5y agoMobile IPs will be a problem until the entire Internet is IPv6. The issue is that there are not enough IPv4 addresses for everyone to get their own IP every time their phone connects to the Internet. So the mobile networks use one IP for many handsets. Block the IP, block dozens of different (innocent) people. Once we're all on IPv6 we can go back to blocking IPs. But then IPv6 creates its own problems.
- greeklish 5y agoHere's a good resource about web scrapping: https://bot.incolumitas.com/#:~:text=more%20sources%2Finformation https://bot.incolumitas.com/#:~:text=more%20sources%2Finform...
- bsamuels 5y ago> I need to make a general remark to people who are evaluating (and/or) planning to introduce anti-bot software on their websites. Anti-bot software is nonsense. Its snake oil sold to people without technical knowledge for heavy bucks. If this guy got to experience how systemically bad the credential stuffing problem is, he'd probably take down the whole repository. None of these anti-bot providers give a shit about invading your privacy, tracking your every movements, or whatever other power fantasy that can be imagined. Nobody pays those vendors $10m/year to frustrate web crawler enthusiasts, they do it to stop credential stuffing.
- nextaccountic 5y ago> Nobody pays those vendors $10m/year to frustrate web crawler enthusiasts, they do it to stop credential stuffing. I don't know about $10m/year, but many sites block bots just because they don't want competitors to access publicly available data. Which is bullshit.
- chucksmash 5y agoThe credential stuffing wiki page didn't exist the last time I thought about invalid traffic so I'm pretty out of date. How is there not an equilibrium here that cuts off credential stuffers? I'd naively imagine the residential IP providers have some measure of bad actors they themselves use to determine if a client is worth it, and that someone getting all your IPs blacklisted would get dropped pretty quickly.
- judge2020 5y agoIn reality residential US ISPs don't really care if their users are getting a sub-par experience since they're often the only fast/fiber provider in the area of their customers, meaning customers have no way to switch. Plus, when a website doesn't work, unless the page itself calls out the ISP (which they never do), customers will think it's an issue with the website and won't possibly attribute blame to their ISP until they're deep in forum threads with people telling them "it's probably your ISP not doing anything about bad customers" - the amount of users going so far to learn this information, then accepting it, is extremely low.
- IceWreck 5y agoHalf of the short-links to cutt.ly aren't working. Why use short links in markdown ?
- al2o3cr 5y agoYou use this software at your own risk. Some of them contain malwares just fyi LOL why post LINKS to them then? Flat-out irresponsible... you build a tool to automate social media accounts to manage ads more efficiently If by "manage" you mean "commit click fraud"
- adinosaur123 5y agoAre there any court cases that provide precedence regarding the legality of web scraping? I'm currently looking for ways to get real estate listings in a particular area and apparently the only real solution is the scrape the few big online listing sites.
- adanto6840 5y agoI was involved in a scraping-related case, though in my situation we were scraping public domain data/facts/public domain media. Email me if you'd like additional info. :) More related to the submission content -- at the time we used rotating proxies, both in-house & external (ProxyMesh - still exists & only good things to say about it); they allowed us to "pin" multiple requests to an IP or to fetch a new IP, etc...
- Grimm1 5y agohttps://en.m.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn https://en.m.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn That’s one of the bigger ones. Unfortunately recent events means scraping is still a gray area.
- amelius 5y agoLegal gray areas are perfect for growth hacking. Just look at Uber and AirBnb.
- Grimm1 5y agoAnd even bigger growth hack for a lot of companies would be having scraping protected by law.
- omgwtfbyobbq 5y agoDo you mean this case? https://en.m.wikipedia.org/wiki/Van_Buren_v._United_States https://en.m.wikipedia.org/wiki/Van_Buren_v._United_States I think it only applies to systems that aren't available to the general public, which in this case was the GCIC. Anything that is available to the public, even if it requires some sort of registration, would I think be legal to scrape. YMMV though.
- curun1r 5y agoThere’s one technique that can be very useful in some circumstances that isn’t mentioned. Put simply, some sites try to block all bots except for those from the major search engines. They don’t want their content scraped, but they want the traffic that comes from search. In those cases, it’s often possible to scrape the search engines instead using specialized queries designed to get the content you want into the blurb for each search result. This kind of indirect scraping can be useful for getting almost all the information you want from sites like LinkedIn that do aggressive scraping detection.
- amelius 5y agoBut won't the search engines block you after some limit has been reached?
- curun1r 5y agoEventually, but they’re not very aggressive when it comes to bot detection. Simple IP rotation usually works.
- marginalia_nu 5y agoI ended up having to put my search engine behind a CDN and CAPTCHA every new IP because I got something like 30,000 search requests like this _per_hour_ from a botnet. They seem to have backed off now but they were really quite persistent for a while.
- tomrod 5y agoSome.
- deleted 5y ago[deleted]
- janmo 5y agoOr, you can spoof the google bot or Bing bot user agent and try to scrape the site that way.
- rfraile 5y agoDatadome, PerimeterX, anyone tried ine if them?
- cmauniada 5y agoI’ve tried skirting datadome but generally you can just get around it by rotating ips, apparently there is a way to de-obfuscate their apps (apps that use datadome services) to retrieve datadome cookies but I haven’t been bothered to check it out yet.
- peterburkimsher 5y agoCouchSurfing uses PerimeterX for profiles.
- lavezzi 5y agoWalmart uses PX and it's pretty easy to bypass.
- ev1 5y agoThey also silently load Threatmetrix now under a walmart domain CNAME.
- rp1 5y agoIt's very easy to install Chrome on a linux box and launch it with a whitelisted extension. You can run Xorg using the dummy driver and get a full Chrome instance (i.e. not headless). You can even enable the DevTools API programmatically. I don't see how this would be detectable, and probably a lot safer than downloading a random browser package from an unknown developer.
- xiamx 5y agoTry your technique on a few of these fingerprint testing sites https://github.com/niespodd/browser-fingerprinting#fingerprint-test-pages https://github.com/niespodd/browser-fingerprinting#fingerpri... I'm pretty sure it's quite detectible
- rp1 5y agoHmm maybe I will if I have time. We've been using this technique for user-initiated scraping. The only issue we've run in to is we get rate-limited by IP sometimes. Changing the IP has solved the problem each time.
- Lukabuz 5y agoIf I am correct in assuming the parent is talking about puppeteer, there is a plugin[1] that claims to evade most of the methods used to detect headless browsers. I have used it recently for just that purpose, and I can say that it worked wifh minimal setup and configuration for my usecase, but I guess depending on the detection mechanisms youre evading YMMV. The creator of that plugin does mention it is very much a cat and mouse game, just like most of the “scraping industry” https://www.npmjs.com/package/puppeteer-extra-plugin-stealth https://www.npmjs.com/package/puppeteer-extra-plugin-stealth
- marginalia_nu 5y agoIf someone is signalling to you you that they do not want your bot on their site, then maybe respect that? Trying to circumvent it is besides being legally questionable, a serious pain in the ass for the site owner and makes websites more prone to attempt to block bots in general. Also, in my experience, most websites that block your bot, block your bot because your bot is too aggressive, or because you are fetching some resource that is expensive that bots in general refuse to lay off. Bots with seconds between the requests rarely get blocked even by CDNs.
- hbgl 5y agoTechnically it should be illegal to scrape websites without the consent of the server owner because it would be a violation of their property rights. That being said, I think that in reality it would be pretty difficult to get courts to agree with you and then to enforce it.
- soheil 5y agoGoogle can access any site without being blocked. They dominate the search space and give little incentive for site owners to allow other bots. I'd say bypassing these measures is fair game while there is a monopoly in search space. We don't want a web that only Google can access. By the way great work on Marginalia search engine, I love it.
- marginalia_nu 5y agoI've honestly not had much problem at all crawling the web as an indie search engine operator. If you want to get past CloudFlare you can register your bot fingerprint with them. A small number of sites has blocked my crawler , but that's almost always been my own fault, and happened a few instances when the crawler was misbehaving and actually fetching too aggressively (or repeatedly). In every case just sending an email to the site explaining what happened and humbly asking for a second chance been enough to be allowed back in. Most website owners don't seem to mind small search engines at all, what they don't want is scrapers that aggressively scrape their entire site 10 times a day, ignoring robots.txt, and being a general nuisance.
- nocturnial 5y agoI knew there was a reason why I used client certificates and alternate ports. Why is it so difficult to just respect robots.txt? Maybe there's an idea for a browser plugin that determines if you can easily scrape the data or not. If not, then the website is blocked and then traffic will drop. I know this is a naive idea...
- remram 5y agoI don't understand what you recommend. Who would drop the traffic?
- ChuckMcM 5y agoI am always amazed when otherwise intelligent people assert without data that the marginal cost of serving web traffic to scrapers/bots is zero. It is kind of like people who say "Why don't they put more fuel in the rocket so it can get all the way into orbit with just one stage?" It sounds great but it is a completely ignorant thing to say.
- ohyeshedid 5y agoSeemingly, most of those people don't have a realistic concept of scale.
- Aperocky 5y agoAnyone who has a minor website knows that majority of the traffic are bot. Imagine if the goal is images and videos, now you've got yourself some heavy duty scraper that could cost the website owner lots of data fees.
- kodah 5y agoWhen I worked in e-commerce as a SRE, bots were doing two things: - trying to disrupt business processes (eg: false referral listings, gift card scams, etc) - trying to disrupt systems I'm sure there are folks who use bots and scrapers for home automation, but these users generate marginal traffic in comparison. The real cost, aside from successfully achieving the points above, is the bandwidth and hardware costs that become overhead. Bots are usually coded with retry mechanisms and ways to change connection criteria on subsequent retries.
- peterburkimsher 5y ago2 of my social media accounts have fallen victim to bot detection, despite not using scripts. There are other websites for which I have used scripts, and sometimes ran into CAPTCHA restrictions, but was able to adjust the rate to stay within limits. CouchSurfing blocked me after I manually searched for the number of active hosts in each country (191 searches), and posted the results on Facebook. Basically I questioned their claim that they have 15 million users - although that may be their total number of registered accounts, the real number of users is about 350k. They didn't like that I said that (on Facebook) so they banned my CouchSurfing account. They refused to give a reason, but it was a month after gathering the data, so I know that it was retaliation for publication. LinkedIn blocked me 10 days ago, and I'm still trying to appeal to get my account back. A colleague was leaving, and his manager asked me to ask people around the company to sign his leaving card. Rather than go to 197 people directly, I intentionally wanted to target those who could also help with the software language translation project (my actual work). So I read the list of names, cut it down to 70 "international" people, and started searching for their names on Google. Then I clicked on the first result, usually LinkedIn or Facebook. The data was useful, and I was able to find willing volunteers for Malay, Russian, and Brazilian Portuguese! After finding the languages from 55 colleagues over 2 hours, LinkedIn asked for an identity verification: upload a photo of my passport. No problem, I uploaded it. I also sent them a full explanation of what I was doing, why, how it was useful, and a proof of my Google search history. But rather than reactivate my account, LinkedIn have permanently banned me, and will not explain why. "We appreciate the time and effort behind your response to us. However, LinkedIn has reviewed your request to appeal the restriction placed on your account and will be maintaining our original decision. This means that access to the account will remain restricted. We are not at liberty to share any details around investigations, or interpret the terms of service for you." So when the CAPTCHA says "Are you a robot?", I'm really not sure. Like Pinocchio, "I'm a real boy!"
- arp242 5y agoCouchSurfing is just shit, full stop. I love the concept and hosted many people, but the way the company has been run over the last few years is beyond atrocious. It's like AirBnB sent over some people to intentionally run it in to the ground or something. LinkedIn has to deal with a lot of scummy recruiters and scammers; I don't blame them for being very strict.
- billpg 5y agoYou could ask first. The site's robots.txt file might have some information. Put your email address in your User-Agent string so they can get in touch if needed.
- deleted 5y ago[deleted]
- walrus01 5y agoGoogle "residential proxies for sale" if you want to see the weird shady grey market for proxies when you need your traffic to come from things like cablemodem operator ASNs' DHCP pools
- spookthesunset 5y agoWonder what fraction of that traffic is from p0wned IoT refrigerators, smoke detectors or WiFi enabled light bulbs… probably more than anybody cares to admit…
- walrus01 5y agoAlso a lot of people who've been tricked into installing malware on their windows PCs, from shady "VPN" operators and other
- RNCTX 5y agoThey've not been tricked, really. The majority of them want to do something completely benign like see a BBC show in the US, or watch an American football show in the UK, and the one defining feature of capitalism is its many contradictory faces. One corporation wants to arbitrarily limit its customer base, and the other corporation wants to arbitrarily limit what data its customer base can see. In between the two is a space big enough to drive a Mack truck or a lorry through, depending on where you're from... One might justify this until the one corporation merges with the other and then you have a situation where the same corporation wants to do two different things to the same pool of users.
- walrus01 5y agoWhen I say tricked, I mean they have no idea that unknown 3rd parties' grey market traffic is being run through their personal home internet connection. They believe they are just using a VPN to watch BBC.
- 5y ago
- completelylegit 5y ago* Scrape open proxy websites for open proxies, then use those proxies, cycle which proxies you use frequently. * Change your user-agent to a real user-agent, cycle it frequently. * Done.
- Jenk 5y agoIn a previous venture my team successfully circumvented bot detection for a price comparison project simply by using apify.com. Wasn't that expensive, either. We were drilling sites with 500k+ hits per day for months.
- connectsnk 5y agoFor the row "Long-lived sessions after sign-in" the author mentions that this solution is for social media automation i.e. you build a tool to automate social media accounts to manage ads more efficiently. I am curious by what the author means by automating social media accounts to manage ads more efficiently
- namdnay 5y agoClickfraud
- dpryden 5y agoIt always amazes me how people believe they have a right to retrive data from a website. The HTTP protocol calls it a request for a reason: you are asking for data. The server is allowed to say no, for any reason it likes, even a reason you don't agree with. This whole field of scraping and anti-bot technology is an arms race: one side gets better at something, the other side gets better at countering it. An arms race benefits no one but the arms dealers. If we translate this behavior into the real world, it ends up looking like https://xkcd.com/1499 https://xkcd.com/1499
- zarzavat 5y agoBecause often that data is only available through scraping. Nobody wants to scrape, it's messy and fickle and a general pain in the backside. But sometimes the data you need exists only in that form. If you run a website and you have a problem with scrapers, then make all that data available through an API and say what acceptable rate limits are. If cost is an issue, then charge a proportionate fee, my time writing a scraper is worth much more than paying a few dollars for an API. If you just say "No" to everything then you lose all control over the process and the only outcome will be such an arm race.
- kingcharles 5y agoGod. This. The number of times I've spent 2 days of my very expensive time coding a scraper to get data I'll use once, when I would have paid a few dollars just to download it in a text file.
- lavezzi 5y agoThe proxy service recommendations are pretty expensive. Does anyone have alternatives they suggest to keep costs down?
- intricatedetail 5y agoBy the way - is it possible to stop Google bot from scrapping without maintaining a list of IP addresses? Google doesn't publish these and it's not good to run reverse DNS as it slows down legitimate clients. I know you can put a meta tag, but bot still has to make a request to read it. I would like to completely cut off Google from scrapping.
- jrockway 5y agoYou can buy databases of who owns which IP blocks. If you really care but don't want to spend the money, just block the subnet each time you see a Googlebot request. "whois w.x.y.z" returns an entire CIDR, and it seems unlikely to me that Google is scraping from a bunch of disconnected /24s.
- drivebycomment 5y agoJust put robot.txt and block Googlebot from there. Google obeys robot.txt.
- intricatedetail 5y agoNot all Google crawlers obey it. Also if Google already indexed something, only way is to let it crawl again and see meta noindex. It's a mess.
- janmo 5y agoSome how related: https://news.ycombinator.com/item?id=29062027 https://news.ycombinator.com/item?id=29062027
- navels 5y agoI've had a lot of success just with Selenium and this custom version of Chromedriver: https://github.com/ultrafunkamsterdam/undetected-chromedriver https://github.com/ultrafunkamsterdam/undetected-chromedrive...
- welanes 5y agoAnother great resource is incolumitas.com. A list of detection methods are here: https://bot.incolumitas.com/ https://bot.incolumitas.com/ I run a no-code web scraper (https://simplescraper.io https://simplescraper.io) and we test against these. Having scraped million of webpages, I find dynamic CSS selectors a bigger time sink than most anti-scraping tech encountered so far (if your goal is to extract structured data).
- kingcharles 5y agoCan your scraper be used to scrape images? I need to scrape some books from a paywalled site and they are presented a page at a time. The JS code is too complex for me to bother trying to figure out how it creates the unique tokens it applies to every image it displays to avoid a very simple scrape.
- teeray 5y agoNever underestimate the scraping technique of last resort: paying people on Mechanical Turk or equivalent to browse to the site and get the data you want
- softwaredoug 5y agoA lot of web scraping is annoying often because there’s *an explicit API built for the scrapers needs*. Instead of looking for an API, many think to first use web scraping. This in turn puts load and complexity on the user facing web app that must now tell scraper from real users.
- fragmede 5y agoBut if there's an API, then the overall load is the same, no? Or to put it another way, naively, having api.example.com and realpeople.example.com separated out into separate sandboxes seems reasonable, but due to the aforementioned problem, its not. But then it also turns out to be the wrong axis for this anyway, and you need your monitoring to work for you.
- kingcharles 5y agoNo, the load isn't the same because the web page might be a multi-megabyte monster piece of badly-coded HTML that returns only say 10 out of 1,000,000 results and needs to be paged through, where the API might return all the million results in a nice JSON chunk.
- no_time 5y agoUsing the API almost always has more "strings attached". Like you have to register and get an API token or something. Or even pay. If you want people to use your API, don't make it less convinient than scraping the page.
- deleted 5y ago[deleted]
- ufmace 5y agoWhat I really enjoy about this thread is all of the completely different perspectives. Lots of people doing anti-abuse research bemoaning that this stuff exists, and lots of people working against what are from their perspective ham-handed anti-abuse tech blocking legitimate useful automation trading tips on how to do it better. I guess the other sides of those we don't see much. People doing actual black-hat work probably don't post about it on public forums, and most of the over-broad anti-abuse is probably a side effect of taking some anti-abuse tech and blindly applying it to the whole site just because that's simpler, often no tech people may be really involved at all.
- matheusmoreira 5y agoWhen these companies endeavor to stop abuse, they trample all over our freedoms. Suddenly we can't have non-browser user agents anymore. Suddenly we can't root our smartphones anymore. They want nothing to do with us unless it's 100% on their terms with us completely under their control.
- bryan_w 5y agoYeah, but you're not entitled to use their servers. If your use of their servers is something they don't like, their freedom is to blackhole your packets
- matheusmoreira 5y agoIt's always these "take it or leave it" deals with these people, isn't it? This is why adversarial interoperability is the rule today.
- Spivak 5y agoThere seems to be wildly different perspectives on "bad actors means we can't have nice things" -- one group says that this is a fact of life, and the other says that this is an affront to freedom. A non-tech example is I've had guys on Tinder get legit angry at me for insisting that our first few dates have to be in public places where we drive separate -- "oh so you think I'm some creepy stalker?" And like I am totally empathetic to their hurt because I'm sure that they know they're good but I don't and there's now way for me to tell in advance. Malicious actors don't exactly announce themselves and actively try to hide their intent. The solution to the automation problem is to do what most companies do and have registered API integrations.
- nuker 5y agoWill I scrape faster with RTX 3080 Ti?
- shapefrog 5y agoAbsolutely, but to get one you will have to have RTX 3090 scraping speeds.
- 0xlwj 5y agoPretty useful crash course on what is out there in the web scraping universe
- egberts1 5y agoA couple of things for unblockable scraping 1. plenty of VPS with many IP addresses (this is easier with IPv6 subnet) 2. HTTP header rearranging 3. Fuzzing user-agent 4. Pseudo-PKBOE algorithm 5. office hours, break-time, lunch-time activity emulation 6. ???? 7. profit I am looking at you, SSH port bashers.
- lifeisstillgood 5y agoWhat if we solved it by replacing passwords with client HSMs?
- hk1337 5y agoNot to forget the most important rule, don't be an asshole to the site hosting the content.
- kseifried 5y agoTrying to stop credential stuffing by blocking bots will not work, and can often severely impact people depending on assistive technologies. I think a better solution is to implement 2FA/MFA (even bad 2FA/MFA like SMS or email will block the mass attacks, for people worried about targeted attacks let them use a token or software token app) or SSO (e.g. sign in with Google/Microsoft/Facebook/Linkedin/Twitter who can generally do a better job securing accounts than some random website). SSO is also a lot less hassle in the long term that 2FA/MFA for most users (major note: public use computers, but that's a tough problem to solve security wise, no matter what). Better account security is, well, better, regardless of the bot/credential stuffing/etc problem.
- firerfly 5y agoplivo.com is good at anti-bot, i tried many method and some residential proxys . there still blocked me out .
- alam2000 5y ago
- kinderjaje 5y agoI am running a no-code web automation and data extraction tool called https://automatio.co https://automatio.co. And from my experience most of the time when using quality residential proxies you will be fine. But that comes at cost since they are way expensive then data center proxies. But for some websites, even residential ips doesn't let you pass. I noticed there is like a premium reCaptcha service, which just work differently then standard one and not let you pass. It's mostly shown with a Cloud flare anti bot page.