18 ms·
Scrape like the big boys
- abc03 5y agoI scrap government sites a lot as they don't provide apis. For mobile proxies, I use the proxidize dongles and mobinet.io (free, with Android devices). As stated in the article, with cgNAT it's basically impossible to block them as in my case, half the country couldn't access the sites anymore (if you place them in several locations and use one carrier each there).
- exhilaration 5y agoWow, this is super interesting: https://proxidize.com/ https://proxidize.com/ https://mobinet.io/ https://mobinet.io/ I feel like I'm getting a glimpse into the dark underbelly of the web.
- rdtwo 5y agoIs it just one ip per dongle at a time? Or can you have multiple ips on the same device.
- KuhlMensch 5y agoDoing a bit of low-stakes monitoring of webpages lately. It started (as I'm assuming it often does) with right-clicking a network request in Chrome and selecting "copy as curl" Then graduated to JavaScript for surrounding logic e.g. data transformation I had assumed I'd quickly give up and move to a headless browser, BUT I can't bring myself to move away from tiny CPU utilization of curl. Throwing together a "plugin" probably takes me less than 20 minutes normally. I'll probably have a look at using prowl to ping my phone. And if I get more serious I'll look at auto authenticate options on npm. But I'm not sure if the overhead of maintaining a bunch of spoofy requests will be worth it.
- ebbp 5y agoHaving spent a week battling a particularly inconsiderate scraping attempt, I’m quite unsurprised by the juvenile tone and fairly glib approach to the ethics of bots/scraping presented by the piece. For the site I work for, about 20-30% of our monthly hosting costs go towards servicing bot/scraping traffic. We’ve generally priced this into the cost of doing business, as we’ve prioritised making our site as freely accessible as possible. But after this week, where some amateur did real damage to us with a ham-fisted attempt to scrape too much too quickly, we’re forced to degrade the experience for ALL users by introducing captchas and other techniques we’d really rather not.
- jtdev 5y agoConsidering the demand for your content, why haven’t you created and provided an API? Maybe you could monetize?
- chewmieser 5y agoLike everyone and their brother has a web spider. And some of them are VERY badly designed. We block them when they use too many resources, although we'd rather just let them be. Can't speak for the op but we have APIs and move the ones scraping and reselling our content to APIs. The majority are just a worthless suck on resources though.
- ebbp 5y agoWe do offer an API - the scrapers are trying to circumvent using that, presumably.
- purerandomness 5y agoWhy do you think are they trying to circumvent it? Does your API provide all the information that can be found on the site, or are they scraping because the API is incomplete? We've once had to scrape Amazon product pages because they have a lot of API endpoints, but those didn't contain the data we needed.
- IceWreck 5y agoThe author says proxys are expensive and then proceeds to spend a shitton of money buying all that hardware.
- palijer 5y agoThat was not the authors main argument against proxies, that was just an additional point. You ignored the primary argument in your judgment. >>Because I could not fully trust the other customers with whom I shared the proxy bandwidth. What if I share proxy servers with criminals that do more malicious stuff than the somewhat innocent SERP scraping?
- RandomThrow321 5y agoCan they not call out a secondary point?
- pnt12 5y agoSure but nitpicking does not lead to productive discussions.
- incolumitas 5y ago4G proxies are just soo much better than so called "residential" or straight datacenter proxies. It makes sense to create your own 4G proxy farm if you conduct business in that area. With only 10 dongles and 10 dataplans, you can have a lot of IP addresses that are extremely hard to block. It's an one time investment, paying proxy providers is a fixed cost.
- bsder 5y agoWhere do you get 4G dongles that don't suck nowadays? We tried to get some, but all of the ones we could get were various levels of broken or unsupported.
- joekrill 5y agoA little pet-peeve I have is when an obscure(ish) acronym is used and never defined. Is SERP a well-known acronym? Perhaps this is a niche blog and I'm not the intended audience.
- tptacek 5y agoYes; a SERP is a Google search result page. It's the most important acronym in SEO.
- nomdep 5y agoI don’t remember never ever hearing it and I’ve been in the industry for some time
- hollerith 5y agoHuh. I've never been in the industry, but noticed "SERP" at least 15, maybe 20, years ago and have remembered it since. (If I were writing something to be published, though, I would write "search-engine results page" instead of "SERP".)
- weird-eye-issue 5y agoYou've been in the SEO industry for some time and never heard SERP?
- nomdep 5y agoThe web programming industry, not the SEO industry
- weird-eye-issue 5y agoThere is actually only very little overlap between SEO and web development. Web developers should know some basic technical SEO but you'd be surprised how much knowledge is in the SEO industry that doesn't overlap at all with development. So, it's not too surprising you don't know what SERP is. Web devs might think SEO just means fast load times, proper markup, meta tags, etc but that's only the very surface
- biosed 5y agoI used to lead Sys Eng for a FTSE 100 company. Our data was valuable but only for a short amount of time. We were constantly scraped which cost us in hosting etc. We even seen competitors use our figures (good ones used it to offset their prices, bad ones just used it straight). As the article suggest, we couldn't block mobile operator IPs, some had over 100k customers behind them. Forcing the users to login did little as the scrapers just created accounts. We had a few approaches that minimised the scraping: Rate Limiting by login, Limiting data to know workflows ... But our most fruitful effort was when we removed limits and started giving "bad" data. By bad I mean alter the price up or down by a small percentage. This hit them in the pocket but again, wasn't a golden bullet. If the customer made a transaction on the altered figure we we informed them and took it at the correct price. It's a cool problem to tackle but it is just an arms race.
- ransom1538 5y agoI love the honey pot approach. Put tons of valued hrefs on the page that are invisible (css) that the scrapper would find. Then just rate limit that ip address and randomize the data coming back. Profit.
- histriosum 5y agoI think this falls into the "arms race" trap, though. If you can make an href invisible via CSS, then the scraper can certainly be written to understand CSS, and thus filter out the invisible hrefs..
- rootusrootus 5y agoI know a guy at Nike that had to deal with a similar problem. As I recall, they basically gave in -- instead of trying to fight the scrapers, they built them an API so they'd quit trashing the performance of the retail site with all the scraping.
- matheusmoreira 5y agoYes. That's exactly what everyone should do.
- devops000 5y agoCould you share your code for AWS lambda and puppetter? It’s definitely interesting for other websites
- incolumitas 5y agoSure. https://github.com/NikolaiT/Crawling-Infrastructure https://github.com/NikolaiT/Crawling-Infrastructure And here I am writing about it (but its quite old): https://incolumitas.com/2019/08/31/web-scraping-puppeteer-aws-lambda/ https://incolumitas.com/2019/08/31/web-scraping-puppeteer-aw...
- neals 5y agoIn a particularly hard to scrape website, using some kind of bot protection that I just couldn't reliably get working (if anybody wants to know what that was exactly, I'll go and check it) I now have a small Intel NUC running with firefox that listens to a local server and uses Temper Monkey to perform commands. Works like a charm and I can actualy see what it's doing and where it's going wrong. (though it's not scalable, of course) We use it for data-entry on a government website. A human would average around 10 minutes of clicking and typing, where the bot takes maybe 10 seconds. Last year we did 12000 entries. Good bot.
- funnyflamigo 5y agoI'm curious what bot protection it was? It couldn't have been trying too hard unless you were employing multiple anti-fingerprinting techniques, I'm assuming you used firefox's built in anti-fingerprinting?
- nkozyra 5y agoYou can use chromium/chrome/cdp and turn headless off and see the same thing.
- kerokerokero 5y agoThanks for the share. Great stuff. I used to scrape websites to generate content for higher SERPs. Ended up going into the adult industry lols. (https://javfilms.net https://javfilms.net)
- anon9001 5y agoNeat! I've run across your site organically :P I've always wondered, and since you're right here... how do sites like this make money? It looks like you're probably crawling all the JAV vendors, finding free clips of today's releases, embedding them in your own site to draw traffic, and making money with affiliate links to buy the full content? Am I missing anything? It seems hard to believe you'd get enough affiliate signups to make it worthwhile. I can imagine your site as being a few hours a year of script maintenance and a money printer, or a 40hr/week SEO job with 1000s of similar sites across the adult industry. I'd love to know anything you're willing to share about how the business works.
- InvOfSmallC 5y agoWhere I was working we stopped caring about ips browser etc because it was just a race. What we did was analyzing behaviour of clicks and acted on that. When we recognized it we went on serving a fake page. It cuts down a little bit of costs because it was static pages. In general it took a lot of time for them to discover the pattern and it was way more manageable for us.
- ahofmann 5y agoWe did the same and the bot developers wrote bots that acted like humans. It took them not very long to find out.
- wilg 5y agoNot the same kind of scraping, but does anyone have thoughts/resources/best practices for doing link previews (like Twitter/iMessage/Facebook)?
- kall 5y agoYou shouldn‘t really need to do any scraping tricks to get that, because it‘s data the websites (usually) want to give to bots. Or are people getting bot block screens from Cloudflare et all for that basic action these days? It should be a matter of a simple GET request to fetch plain html and parse the OpenGraph meta tags out if that. There are many open source libraries to do that for you depending on your language. If bot blocks really are a problem, a SaaS solution like Microlink could probably do it for you.
- wilg 5y agoBot blocks are definitely an issue for certain sites, I've implemented it that way currently. Microlink is a good tip, thanks!
- mrg3_2013 5y agowow! That was an interesting read.
- hall0ween 5y agoBasic question, how does one profit from scraping data and what kinda data? Taking a stab at answering it: you scrape the data and build a business around selling it. Stock prices? But that's boring, plus how many others are doing it? I bet a lot.
- 323 5y agoThese are scraping artificially limited releases clothes/shoes. You buy a shoe at $100 and immediately sell it at $1000. Artificial scarcity - every week you release a "limited edition item", but if you do the math, it's not limited edition at all if you integrate over a year.
- ushtaritk421 5y agoHere's a project that's been in the news recently that relies heavily on scraped data. http://www.thebillionpricesproject.com/ http://www.thebillionpricesproject.com/
- throw1234651234 5y ago1. Be job site. 2. Have employees that cost money call facilities and get job listings. 3. Establishing relationships with facilities to list jobs. 4. Buy job listings from 3rd parties. 5. List them for free hoping to make margin. 6. Scraper steals all jobs, lags site, and gets value of hard work for free.
- hall0ween 5y agoahh thanks
- ushtaritk421 5y agoAnything you might look up or keep track of online that helps your business is probably being scraped by someone who is using it themselves or selling access to the curated data set. Prices (are yours high, low compared to competition?), reviews, locations of physical stores, search result placement (where does your widget show up when someone searches "widget" on your site?), just to name a few use cases.
- DeathArrow 5y agoYou can put some wasm crypto mining code and at least profit from bots. :D
- max002 5y agoIts easy to detect chrome headless so scraping with it is not really how "big" boys do it :D the only scrapers/bots that are really hard to detect are the ones running and controlling real browser and not chromium. I do a lot od research aggainst abitbot systems, some times is friday night. If you spend each one in pub it doesnt mean your normal.
- beaugunderson 5y agopuppeteer-extra and undetected-chromedriver beg to differ :)
- max002 5y agoNot really, i did test it (and use it for some cases), but there are still sites that detects it. I can and anyone who can check webgl renderer name, though this can be done by faking driver name but thats just one of many ways:) Its ongoing fight. If you dont move your mouse or type faster than 95% of my portal users i can detect you with js script written in under 1 minute.
- incolumitas 5y ago:D
- throwaway984393 5y agoIf you want to avoid bot detection, learn how bot detection work. A lot of commercial "webapp firewalls" and the like actually have minimum requirements before they flag certain traffic as a botnet; stay below those limits and you can keep hammering away. Sometimes those limits are quite high. In the past we've had the most success defeating bots by just finding stupid tricks to use against them. Identify the traffic, identify anything that is correlated with the botnet traffic, and throw a monkey wrench at it. They're only using one User Agent? Return fake results. 90% of the botnet traffic is coming from one network source (country/region/etc)? Cause "random" network delays and timeouts. They still won't quit? During attacks, redirect to captchas for specific pages. During active attacks this is enough to take them out for days to weeks while they figure it out and work around it.
- chrisMyzel 5y agoWe are seeing a lot of bot traffic too but chose to accept it as reality. We are aware if thousands of bots create unpredictable cost surges that there is something wrong with our product, it should not create such heavy loads to our servers in the first place to fulfil it's mission. I believe the future will make us more free by using more bot / AI technology since who wants to spend their whole day in front of a computer and research information if machines can do the job just fine?