4 ms·
Isn't it impossible to win the game of blocking headless browsers? What's stopping someone from creating an API that opens up a real browser, uses a real (or v
by tabeth 9y ago
Isn't it impossible to win the game of blocking headless browsers?
What's stopping someone from creating an API that opens up a real browser, uses a real (or virtual) keyboard, types in/clicks the real address, etc. then proceeds to use computer vision to scrape the information from the page without touching the DOM?
- ThrustVectoring 9y agoThere's a similar "analog hole" for video DRM, too.
- calebegg 9y agohttps://en.wikipedia.org/wiki/Analog_hole https://en.wikipedia.org/wiki/Analog_hole
- mr_toad 9y agoI wonder how long it will be before someone comes up with the idea of using iPhone style facial recognition to tell whether a human is looking at the TV/Monitor or not.
- ThrustVectoring 9y agoOh please no, exposing those sorts of APIs will quickly be utilized by ad-tech guys to make interstitial video ads that don't go away until you finish watching them.
- nicklaf 9y agoBlack Mirror S01E02
- dsjoerg 9y agoYou are in principle correct, but in practice you need to account for the side channels of information as well -- does the mouse and keyboard behave like a human or a robot? Are there thousands upon thousands of sessions coming from the same IP address? The cat and mouse game happens at every level, not just the DOM/browser-detection level.
- patcheudor 9y ago>Are there thousands upon thousands of sessions coming from the same IP address? As someone who routinely works behind proxies, I can sympathize strongly with this statement: "The one thing that I was really trying to get across in writing that is that blocking site visitors based on browser fingerprinting is an extremely user-hostile practice."
- nicklaf 9y agoSo record actual user input data and generate similar input patterns stochastically. That said if you try to scale this up beyond what a reasonable, normal user world do in one sitting, you are bound to stand out. Although that said, I find that I trigger such rate-limiting mechanisms already as a human just when searching Google as a human being and clicking through every last search result page.
- averagewall 9y agoYou'd have to scrape slowly to mimic a real slow user. Maybe at that point you'd be cheaper to get Mechanical Turk to do it. That should solve IP rate limiting, captchas, and just about everything except the endless arms race. Why are so many people going directly to these same-formatted internal URLs without clicking through from random other places? So the site can change the internal URLs and break it all over again.
- figgis 9y ago>You'd have to scrape slowly to mimic a real slow user. Sure, but that's easily mitigated by running multiple scrapers as different users.. You don't need to get all the data from a single scrape.
- toomuchtodo 9y agoYou'd use a browser extension, scoped to requests of sites you're interested in, and stream your data back to your infrastructure for processing. You're limited only by your install base and your ingest infrastructure. Recap [1] does this to extract PACER court documents that are public domain, but access is restricted due to draconian public policy. [1] https://free.law/recap/ https://free.law/recap/
- bryanrasmussen 9y agowell headless browsers exist because they are less expensive to automate than real browsers, adding in the computer vision and the scraper just added a lot of expense.
- roywiggins 9y agoRate-limiting followed by CAPTCHAs seems to be the usual strategy. I think Google claims to try to detect humans by parsing out their mouse movements and scroll events.
- nicklaf 9y agoAnd I can attest that they often presume that I'm a robot. At this point it would be easier for me to write an alternative frontend to Google search (or just use duckduckgo), but it was be amusing to think that I might evade this by writing a script to simulate mouse movements to appear less robotic.
- mehrdadn 9y ago> And I can attest that they often presume that I'm a robot. In my experience this occurs when you are either doing this too much, or you are not accepting their cookies when logged in. (I don't recall the behavior when logged out.)
- nicklaf 9y agoFor the record: I was logged out / incognito. Actually, for whatever reason, this also often seems to cause reCapcha to essentially hellban me and just keep asking for me to solve an endless series of capchas. :/
- srirachahot 9y agoOne simple reason is resource cost. Having a non headless browser is more expensive. Therefore at scale you are wasting resources.
- adrianmonk 9y agoYes, it's more expensive than a headless browser. Impossible to say, though, whether it's more expensive than can be justified in achieving an objective until that objective is known.
- temuze 9y agoOr just recompile Chromium with a few changes ;)