16 ms·
It is not possible to detect and block Chrome headless
- fixermark 9y agoI'm not sure why one wants to bother to do this. With tools like Sikuli script (sikuli.org) already around for ages, automating a headed browser isn't rocket science. So the best-case scenario for detecting headless browsers is "The bad guys just use headed browsers and another automation solution."
- Dolores12 9y agoLooks like great tool, never heard of sikuli before. Thanks for the tip!
- fixermark 9y agoI name-drop it every chance I get ;) We used it to automate the integration tests for a game engine at a previous company; worked great, because it allowed us to fire events into the engine itself based upon the actual rendered pixels (Sikuli supports varying levels of fuzzy image detection for event targets).
- deleted 9y ago[deleted]
- heipei 9y agoInteresting follow-up (again). It will be very interesting to see where attempts to detect headless browser will first appear in the wild. Once we know that and the prevalence, we can make a judgement call on how much effort to put into anti-detection techniques. It's an arms race for sure, but once you know your target you can evaluate whether you even have to put up the effort to defeat a non-existent adversary.
- draw_down 9y agoSome people claim it’s rabbit season. Others contend that in fact it is duck season.
- deleted 9y ago[deleted]
- nukeop 9y agoGood, the less effective various spying techniques are, and the easier they are to throw off, the better the internet is for its users. I don't want any website owners to know what device, browser, or other program, I use to access their site, and they have no business knowing that. I like it being a piece of information I can supply voluntarily for my own purposes, and I get the heebie jeebies every time I read about a new shady fingerprinting technique that exploits some new, previously unexplored quirk of web technologies.
- zzzcpan 9y agoThis incentivizes more aggressive fingerprinting, not the other way around. Too bad people don't realize it.
- ladzoppelin 9y agoBrowser fingerprinting, I almost forgot. Non aggressive and impossible to stop.
- cortesoft 9y agoThe conclusion that this makes spying techniques more difficult is not at ALL what this article is saying. It is just saying that no fingerprinting is going to be able to distinguish between headless and not headless, but that is because there are too MANY variations, not because the variations are hard to detect. Nothing in this article gives any instruction or guidance on how to prevent your browser from being fingerprinted and tracked.
- jimrandomh 9y agoSites detecting headless browsers vs headless browsers trying not to be detected by sites, is an arms race that's been going on for a long time. The problem is that, if you're trying to detect headless browsers in order to stop scraping, you're stepping into an arms race that's being played very, very far above your level. The main context in which Javascript tries to detect whether it's being run headless is when malware is trying to evade behavioral fingerprinting, by behaving nicely inside a scanning environment and badly inside real browsers. The main context in which a headless browser tries to make itself indistinguishable from a real user's web browser is when it's trying to stop malware from doing that. Scrapers can piggyback on the latter effort, but scraper-detectors can't really piggyback on the former. So this very strongly favors the scrapers.
- troels 9y agoIn my experience, the most effective counter measure to scraping is not to block, but rather to poison the well. When you detect a scraper - through what ever means - you don't block it, as that would tip it off that you are on to it. Instead you begin feeding plausible, but wrong data (like, add a random number to price). This will usually cause much more damage to the scraper than blocking would. Depending on your industry etc., it may be viable to take the legal route. If you suspect who the scraper is, you can deliberately plant a 'trap street', then look for it at your suspect. If it shows up, let lose the lawyers. Of course, the very best solution is to not care about being scraped. If your problem is the load that is caused to your site, provide an api instead and make it easily discoverable.
- zzzcpan 9y ago"That’s when it becomes impossible. You can come up with whatever tests you want, but any dedicated web scraper can easily get around them." As long as the logic is hidden from the scrapers, i.e. not running in a web browser, scrapers are at a disadvantage. They don't have the data about the users that websites have. And even something as simple as Accept-Language header associated with an IP subnet is a data point that can be used to protect against scraping. There are a lot more data points though and more aggressive fingerprinting can effectively destroy scraping.
- hartleybrody 9y agoThe counterpoint (and what this article mentioned in several places) is that the "more aggressive fingerprinting" techniques can have high false positive rates, and then the legitimate users of your site are going to end up thinking it's broken or the data never loads because they're inadvertently triggering your fingerprinting. It feels pretty ridiculous to tell a user "you can't use our site because your system language settings don't match our IP list for that area." Sidestepping the fact that assuming you can tell someone's language from the geo mapping of their IP address is already pretty problematic.[0] [0]: https://medium.com/@kristopolous/stop-guessing-languages-based-on-ip-address-3862464b97c7 https://medium.com/@kristopolous/stop-guessing-languages-bas...
- dsp1234 9y agoIf I can navigate to it using a normal Chrome instance, then a headless instance is going to have all of the same information (accept-language) as far as the server can detect. Chrome headless is Chrome and makes network calls in the exact same way. That's why client side headless detection even exists. So the only way in which what you suggest would work is if the system is so aggressive that it actually blocks a normal Chrome instance, which is more hostile that most systems are. But this allows a user to change settings until they do get a correct response back, and then just have the headless browser use those settings.
- kbenson 9y agoAll the passive techniques are much harder to reasons about, but much easier to match. You just look at the complete request/response headers, make sure you match them, and have some good sources to request from. Much harder is stuff like Distil's script injection, where they transparently inject script tags that do fingerprinting, and they obfuscate the code that does so annoyingly (it's not really hard to reverse, just time consuming and and annoying). They pair this with being a bit more user friendly by redirecting you to a CAPTCHA page if your fingerprinting hits some threshold which if you answer redirects you to the page you wanted, so users experience and inconvenience if there's a false positive, but still get access to what they wanted. I was able to get around most of the passive fairly easily with Perl and LWP, and even the active stuff and CAPTCHA redirects (cookie_jar all serialized to DB so I could store request and represent it to a user to answer), but once they started tweaking their fingerprint script ever couple months/weeks that's when the equation shifted. Distil, as a solutions provider, gets to amortize their changes across all their customers, while I would have to spend the time de-obfuscating it. They could just assign a person to change it once a week and they would effectively halve my time to get any real work done, so without a collective effort of some sorts to combat them, I saw the writing on the wall. :/ The sad thing is that when we moved to API access, their APIs are hampered to the degree that it actually takes two orders of magnitude more requests each minute for a fraction of the accuracy (I was able to query changes over the last couple minutes previously, and now I have to query the entire item set of a subset of all containers, when there are tens of thousands of containers). :/ Lose lose, since our use case isn't even the main reason the site wanted to block scrapers.
- merb 9y agoif you want to block scrapers, just add rate limiting...
- lr4444lr 9y agoI'll eat you through a proxy network then, unless you want to slow down your legitimate users too.
- thaumaturgy 9y agoI now work for a company that is gathering metadata on the IP address space (in an effort to reduce the amount of abuse that sites and service providers have to deal with). It won't be very long before it'll be possible to identify most of the common proxying networks and block those. Scrapers can respond by setting up something like an ssh tunnel from a residential high speed connection to a remote server (so that scraping run on the remote server appears to come from the residential connection) and using that as a proxy, but it adds another level of effort. (I am generally on the scrapers' side on this and think it's ultimately futile to try to block all scrapers. Website administrators need to accept that anything they put on the internet is public and get over it.)
- Jach 9y agoMaybe companies can instead focus on improving network infrastructure and software architecture to be able to eat any amount of scraper traffic since at some point it must become indistinguishable from a ddos that also needs to be handled...
- tehlike 9y agoIts not about ddos. Think amazon. They probably dont want to reveal all of heir inventory, their price history and so on. Even their api had rate limiting per second, and i doubts it is because of the load
- deleted 9y ago[deleted]
- mixedbit 9y agoIt is impossible to make headless and normal browser send 100% indistinguishable traffic. The timing of the browser requests is influenced by rendering that for the two versions will be always different.
- seanp2k2 9y agoIt's not impossible; you could e.g. profile a real client's timing and introduce delays into the headless version. It's not zero work, but it's very much not impossible if you're sufficiently motivated. Especially recently with e.g. https://hackaday.com/2018/01/06/lowering-javascript-timer-resolution-thwarts-meltdown-and-spectre/ https://hackaday.com/2018/01/06/lowering-javascript-timer-re... , high-precision timers in JS might not be available for all clients for reasons other than ~"they're headless and trying to scrape my site".
- gildas 9y agoThe problem is that you can easily detect that some properties have been overloaded. For example, you can execute Object.getOwnPropertyDescriptor(navigator, "languages") to detect if navigator.languages is a native property or not.
- numbsafari 9y agoCan’t I accomplish the same thing by compiling a modified version of headless chrome?
- gildas 9y agoYes I think it's possible but it's much harder to do, I guess.
- sigotirandolas 9y agoIt's possible to hide that as well, funnily enough, by also overwriting Object.getOwnPropertyDescriptor (and similar tricks). As far as I know, it's theoretically possible to use this trick to completely 'sandbox' some code so that there's no way it can detect certain functions being overwritten (by overwriting all functions such as Object.getOwnPropertyDescriptor, Function.toString, etc. and making them hide the overwritten functions, including themselves). For some more information: http://randomwalker.info/publications/ad-blocking-framework-techniques.pdf http://randomwalker.info/publications/ad-blocking-framework-...
- gildas 9y agoVery interesting article! Thank you.
- beagle3 9y agoThe EME DRM is part of the game for those who really want to block headless. It will arrive, sooner or later.
- OkGoDoIt 9y agoI always assumed DRM would eventually factor into this. I’ve only ever read about it in the context of media, but I’m assuming there’s ways to use it creatively for fingerprinting and blocking scraping as well. Do you have any links with insights into that?
- beagle3 9y agoI am not familiar with the specific details and how they would allow this, but ... if it is for human consumption and not bot consumption, it is enough to render the result into a DRMd h264
- koiz 9y agoGood. It shouldn't. The web should be open the fact that people are still trying to stop this is a joke.
- yorby 9y agotell that to web assembly and friends
- tabeth 9y agoIsn't it impossible to win the game of blocking headless browsers? What's stopping someone from creating an API that opens up a real browser, uses a real (or virtual) keyboard, types in/clicks the real address, etc. then proceeds to use computer vision to scrape the information from the page without touching the DOM?
- ThrustVectoring 9y agoThere's a similar "analog hole" for video DRM, too.
- calebegg 9y agohttps://en.wikipedia.org/wiki/Analog_hole https://en.wikipedia.org/wiki/Analog_hole
- mr_toad 9y agoI wonder how long it will be before someone comes up with the idea of using iPhone style facial recognition to tell whether a human is looking at the TV/Monitor or not.
- ThrustVectoring 9y agoOh please no, exposing those sorts of APIs will quickly be utilized by ad-tech guys to make interstitial video ads that don't go away until you finish watching them.
- nicklaf 9y agoBlack Mirror S01E02
- dsjoerg 9y agoYou are in principle correct, but in practice you need to account for the side channels of information as well -- does the mouse and keyboard behave like a human or a robot? Are there thousands upon thousands of sessions coming from the same IP address? The cat and mouse game happens at every level, not just the DOM/browser-detection level.
- urlgrey 9y agoCrawlers & scrapers that rely on headless browsers like Chrome often initiate playback of video on the pages they access. The company I work for (Mux) has a product that collects user-experience metrics for video playback in browsers & native apps. It's been a non-trivial effort developing a system to identify video views from headless browsers so that we might limit their impact on metrics. Being able to make this differentiation has a real benefit to human users of our customer's websites. My preference would be for headless browsers to not interact with web video or be easily identifiable via request headers, though I doubt either of these things will happen any time soon.
- criddell 9y agoHow are they initiating playback? Are they pressing play, or just triggering auto-play behavior?
- goerz 9y agoVideo should never play unless actively initiated by the user. That would fix the metrics, as the headless browser probably wouldn't initiate the video playback
- bpicolo 9y agoIf it's a video site, I expect the video to play when I land, e.g. youtube. I'm initiating on purpose by browsing
- nybble41 9y agoRegarding YouTube in particular, I tend to open up videos in background tabs for later viewing and find it very annoying that they start playing automatically before I get around to that tab. I did go there to watch the video—eventually. Just not the second that the page finishes loading. YMMV. A persistent setting to enable or disable auto-play would be ideal.
- JeremyBanks 9y ago
- j_s 9y agoThis dicussion is also happening on a counterpoint posted about 9 hours earler, also currently on the front page: It is possible to detect and block Chrome headless | https://news.ycombinator.com/item?id=16175646 https://news.ycombinator.com/item?id=16175646
- pwaai 9y agoSome people seem to have figured out how to detect without relying on fingerprinting the browser. ex. Crunchbase but headless chrome shouldn't be possible to distinguish from a regular chrome browser. The only vector to block scraper is some sort of navigational awareness that deviates from a distribution curve + awareness of IP. but this comes at a great cost to hurting your own real vistors by taxing them with captcha or other annoyances.
- londons_explore 9y agoThats what invisible recaptcha is for. Only the users who compulsively clear cookies ever get bothered by it, and even then all they have to do is click a few photos of cars.
- Sephr 9y agoThe author's navigator.webdriver fix is easily detected, though of course it is fixable with changes to Chrome. This cat and mouse game probably isn't worth pursuing against dedicated adversaries. if (navigator.webdriver || Object.getOwnPropertyDescriptor(navigator, 'webdriver')) { // navigator.webdriver exists or was redefined }
- foob 9y agoThat test actually wouldn't work: > navigator.webdriver true > Object.getOwnPropertyDescriptor(navigator, 'webdriver') undefined As you say though, it's a cat and mouse game and you could always override the behavior of getOwnPropertyDescriptor() if it were used in a test.
- megamindbrian2 9y agoCool, I think the new captchas use mouse entropy that would be an interesting test since remote usually go straight to the pixel point.
- robocat 9y agoMost touchscreens go straight to the pixel point too.
- _o_ 9y agoAll those tests are useless and effective only against script kiddys (which are now like 99.99999% of developers by old standards) and are unable to code anything else but crappy languages like js. For people grown up with web, capable of coding in c/c++ those tests are a joke, I'll just modify the source code to return what is expected and 'game over'. We were reversing drms by dissasembling and patching the binaries - in world of text based protocols and scripts, Idiocracy of todays world is making us invincible.
- icebraining 9y agoWhat does C/C++ have to do with this, when the point of the article is showing that they can be defeated using JS?
- _o_ 9y agoJS is run within c/c++ js engine that can be modified to return you whatever fake results. You can't prevent that. As always, any lock is cheaper to defeat than create.
- icebraining 9y agoNo, you misunderstood; you can defeat the detection using JS. You don't need C at all.
- _o_ 9y agoThis article is a joke, all those methods of "protections" are a joke. What we called "script kiddys" and are now a major amount of so called developers are just underdeveloped lamers who just don't know that the fight is lost in advance. All the methods that you take are useless when you get into situation of scraper run by someone who is able to modify (oh and is able to code in c/c++) and recompile the client side. The world went into Idiocracy so much that methods are beeing invented by people who are so narrow minded that they see the development in a scope of a browser and have a false sense that they "can handle it". Only if the oponent is as narrow minded as they are. Only than. I can modify the source code of chromium, you will get back exactly what you expect from regular user, i am able to scrape fb and linkedin and the only thing they can do is to slow me down (to hide the fact that the code is doing surfing, not human). Stop wasting your time on protection, you are running your inneficient crappy code in insecure environment, the only "attacker" you are safe against is the one who is as clueless as you are. The same moment when you send content to the client, it is game over. You have lost all control. I am sorry for all non-gentle sentences here, but we had developers who were able to decompile asm code and patch it to avoid drms, while now sandboxed idiots are thinking, they are smart. The whole dev. environment became toxic =/ And people are just to stupid to understand how stupid they are =/
- xfer 9y agoHarsh words; but very true that the developers don't realize the chromium is open-source.. maybe they should just jump to the new drm extension, atleast that will challenge the dedicated scrapers.
- mykull 9y agoThis is pretty broad criticism, but you are at least right that there is no sound principle on which to protect a website from "automated" interaction.
- deleted 9y ago[deleted]
- j_coder 9y agoIt is "easy" to block scraping. Make it very costly to scrape: - Render your page using canvas and WebAssembly compiled from C, C++, or Rust. Create your own text rendering function. - Have multiple page layouts - Have multiple compiled versions of your code (change function names, introduce useless code, different implementations of the same function) so it is very difficult reverse engineer, fingerprint and patch. - Try to prevent debugging by monitoring time interval between function calls, compare local time interval with server time interval to detect sandboxes. - Always encrypt data from server using different encryption mechanisms every time. - Hide the decryption key into random locations of your code (use generated multiple versions of the code that gets the key) - Create huge objects in memory and consume a lot of CPU (you may mine some crypto coins) for a brief period of time (10s) on the first visit of the user. Make very expensive for the scrapers to run the servers. Save an encrypted cookie to avoid doing it later. Monitor concurrent requests from the same cookie. The answer is that it is possible but it will cost you a lot.
- tlrobinson 9y agoAll of which is defeated by OCR.
- j_coder 9y agoYes. The idea here is to make you dependent on OCR (you also have to find where is the information as the page design changes) and to waste a lot of your server resources making it very costly to scrape.
- eastendguy 9y agoGood point. OCR powered web scraping is even available out of the box nowadays. https://a9t9.com/kantu/docs/scraping#ocr https://a9t9.com/kantu/docs/scraping#ocr
- j_coder 9y agoIt is not the OCR that is costly. It is the JavaScript execution to render the page so you can do the OCR. You can even increase the JavaScript execution cost if suspicious. You will also have to automate all page variations and the traditional challenges (login, captcha, user behavior fingerprinting, ...) At the end the development time, cost and server cost will kick you out of business if you are too dependent on the information or you start to loose money every time you scrap.
- landryraccoon 9y agoIf you want to detect if a human is visiting your site, open an ad popup with a big close button directly over the content. A human being will always, 100% of the time, immediately close the popup. Automation won't care.
- xienze 9y agoOK, but that is guaranteed to annoy users. Plus, I think you’re underestimating the intelligence of the people writing scrapers — obviously they’re going to visit the site manually and see what appears to be a fingerprinting measure. Then they’ll update the scraper to close that pop up. There are no effective solutions to this problem.
- deleted 9y ago[deleted]
- pbalau 9y agoChromium can work directly with Wayland afaik. Do a "fake" Wayland implementation and Chromium will happily think it's drawing crap
- ed9911 9y agoAs someone who writes web scrapers for a living, I have only come across one site where I have been unable to reliably extract the information we need. If we were more flexible, we would be able to deal with this site too. Defending yourself from scrapers is an arms race you are almost certain to loose.
- deleted 9y ago[deleted]
- odammit 9y agoBlocking crawlers is dead simple: Find a way to build an API for your data that allows you both to make money. Any effort besides that is wasted. Honey pots links? Great my crawler only clicks things that are visible. See capybara. IP thresholds? Great I have burner IPs that hit a good page of yours until I’m blocked (am I time banned, captchad or perma) and then I back that number out across my network of residential IPs (bought through squid, hello or anyone else) and a mix of tor nodes ( I sample your site with that too) to make sure I never approach that number. But then I also geolocate the IP so it’s only crawling during sensible browsing hours for that location. Keystrokes detection? yeah I slow down keystrokes so it looks like Grandma is browsing Mouse detection? looks like Michael j Fox is on your site (that’s an old Dell or Gateway Commercial reference don’t be mad) Poison the well? I get a page from multiple IPs and headless browser combinations on different screen orientations and if I detect odd changes in data I flag that URL for a turk to provide insight/tune the crawler. I keep the screenshot and full payload (css,js,html) that I use over time to do more devious shit like render old versions of your page behind a private nginx server so I can re-extract pieces of data I may have missed. Stop trying to stop the crawling and figure out how to create a revenue stream.
- odammit 9y agoCaveat to my wasted effort comment: Your'e an e-commerce site that has a problem with people buying goods (especially virtual goods, ebooks, gift cards, etc)[1] with stolen credit cards. You need a solution. The hardest thing I've ever had to crawl (as I mention in another comment in this thread) has been linkedin and Facebook. Why? Because I have to be logged in to get the data I want. If you want to stop crawlers you can also put the good stuff behind auth, but you need a solid auth mechanism. You can't just do email verification because I'll use shit like https://www.mailinator.com https://www.mailinator.com to generate a ton of fake emails to sign up for your site. [1] Why virtual goods? You can't stop shipping or track down the person once the card is reported stolen.
- littlestymaar 9y ago> [1] Why virtual goods? You can't stop shipping or track down the person once the card is reported stolen. In the meantime, virtual good are also zero-cost : when you sell a ebook and the transaction is cancelled by the bank, you didn't lose anything, it's not like the buyer was willing to pay anyway.
- aplorbust 9y agoIts trivial to randomise HTTP headers, both the content and the order. There are free and commercial databases of user-agent strings available to any user, the same ones the websites may use. Users can also modify or delete HTTP headers through local proxies, using the same proxy software that many high volume websites use. Sites that rely on redirects to set headers make this even easier. p0f only works with TCP. Could this be another a selling point for alternative congestion controlled reliable transports that are not TCP, e.g. CurveCP? I have prototype "websites" on my local LAN that do not use TCP. The arguments in favor of controlling access to public information through "secret hacker ninja shit" (https://news.ycombinator.com/item?id=16176572 https://news.ycombinator.com/item?id=16176572) are not winning on the www or in the courts. Consider the recent Oracle ruling and the pending LinkedIn HiQ case. If the information is intended to be non-public, then there is no excuse for not using access controls. Anything from basic HTTP authentication to requiring client x509 certificates would suffice for making a believable claim. Detecting headless Chrome and serving fake information, or any other such "secret hacker ninja shit" is not going to suffice as a legitimate access control, whether in practice or in an argument to a reasonable person. The fact is in 2017 websites still cannot even tell what "browser" I am using, let alone what "device" I am using. They still get it wrong every time. Best they can do is make lousy guesses and block indiscriminately. Everything that is not what they want/expect is a "bot", a competitor, an evil villan. Yet they have no idea. Sometimes, assumptions need to be tested.1 1 https://news.ycombinator.com/item?id=16103235 (where developer thought spike in traffic was an "attack")
- baybal2 9y agoFrom my experience in the scene: Bot mill people are very aware of headless browsers being an effortless solution to mimic a browser, but not that efficient. The amount of ram and so a bots spends to do a single click can truly hurt their bottom line. Top tier collectives I heard of use own C/C++ frameworks with hardcoded requests and challenge solvers, and in-depth knowledge of anti-botting anti-fraud techniques used by the opposing force. If DoubleClick finds a brand new performance profiling test, and send it out in the JS code in one in 1000 requests, expect those guys to detect it and crack it within 24 hours. They have no objective of getting through captchas, just having their number of valid clicks in double digits.
- emmelaich 9y agoIt's a very dangerous thing to do for SEO reasons too. I'm sure Google and others have automated user-like crawling which attempts to validate their official Google indexing bot. If the results between the two differ in certain ways you may well get your site buried way down in search results.
- wiz21c 9y agoCould someone tell me why everybody wants to fight against headless browser ? If I want to use such a browser to browse your site, site that you voluntarily show to the public, then it's my problem, my code, not yours. If you want to protect your data so much, then maybe you shouldn't put them on the web first place. (yep, I present things in black and white, but you get the picture) I would also add this : https://www.bitlaw.com/copyright/database.html#Feist https://www.bitlaw.com/copyright/database.html#Feist because it basically says it's hard/pointless to protect data.