9 ms·
Detect and crash Chromium bots
- lifthrasiir 1y agoPreviously on HN: Detecting Noise in Canvas Fingerprinting https://news.ycombinator.com/item?id=43170079 https://news.ycombinator.com/item?id=43170079 The reception was not really positive for the obvious reason at that time.
- chrismorgan 1y agoChecking https://issues.chromium.org/issues/340836884 https://issues.chromium.org/issues/340836884, I’m mildly surprised to find the report just under a year old, with no attention at all (bar a me-too comment after four months), despite having been filed with priority P1, which I understand is supposed to mean “aim to fix it within 30 days”. If it continues to get no attention, I’m curious if it’ll get bumped automatically in five days’ time when it hits one year, given that they do something like that with P2 and P3 bugs, shifting status to Available or something, can’t quite remember. I say only “mildly”, because my experience on Chromium bugs (ones I’ve filed myself, or ones I’ve encountered that others have filed) has never been very good. I’ve found Firefox much better about fixing bugs.
- lillecarl 1y agoI guess it depends on what kind of bug it is, this took 25 years to fix https://news.ycombinator.com/item?id=40431444 https://news.ycombinator.com/item?id=40431444
- Dylan16807 1y agoTo be fair that bug was only P3.
- oefrha 1y ago> The call to page.evaluate just hangs, and the browser dies silently. browser.close() is never reached, which can cause memory leaks over time. Not just memory leaks. Since a couple months ago, if you use Chrome via playwright etc. on macOS, it will deposit a copy of Chrome (more than 1GB) into /private/var/folders/kd/<...>/X/com.google.Chrome.code_sign_clone/, and if you exit without a clean browser.close(), the copy of Chrome will remain there. I noticed after it ate up ~50GB in two days. No idea what's the point of this code sign clone thing, but I had to add --disable-features=MacAppCodeSignClone to all my invocations to prevent it, which is super annoying.
- closewith 1y agoThat's an open bug at the minute, but the one saving grace is that they're APFS clones so don't actually consume disk space.
- oefrha 1y agoInteresting, IIRC I did free up quite a bit of disk space when I removed all the clones, but I also deleted a lot of other stuff that time so I could be mistaken. du(1) being unaware of APFS clones makes it hard to tell.
- omneity 1y ago[flagged]
- randunel 1y agoHow do you deal with the usual CF, akamai and other fingerprinting and blocking you? Or is that the customer's job to figure out?
- omneity 1y agoThank you for the question! It depends on the scale you're operating at. 1. For individual use (or company use but each user is on their device) typically the traffic is drown out in regular user activity since we use the same browser and no particular measure is needed, it just works. We have options for power users. 2. For large scale use, we offer tailored solutions depending on the anti-bot measures encountered. Part of it is to emulate #1. 3. We don't deal with "blackhat bots", so we don't offer support to work around legitimate anti-bot measures such as social spambots etc.
- lyu07282 1y agoIf you don't put significant effort into it, any headless browser from cloud IP ranges will be banned by large parts of the internet. This isn't just about spam bots, you can't even read news articles in many cases. You will have some competition from residential proxies and other custom automation solutions that take care of all of that for their customers.
- omneity 1y agoThanks, that's so true! We learned this the hard way building Monitoro[0] and large data scraping pipelines in the past, so we had the opportunity to build up the required muscle. One thing to note, there are different "tiers" of websites, each requiring different counter-measures. Not everyone is pursuing the high competition websites, and most importantly as we learned in several cases scraping is fully consensual or within the rights of the user. For example: * Many of our users scrape their own websites to send notifications to their discord community. It's a super easy way to create alerts without code. * Sometimes users are locked in their own providers, for example some companies have years of job posting information in their ATS they cannot get out. We do help with that. * Public data websites who are underutilized precisely because the data is difficult to access. We help make that data operational and actionable. We had for example a sailor setup alerts on buoys to stay safe in high waters. A random example[1] 0: https://monitoro.co https://monitoro.co 1: https://wavenet.cefas.co.uk/details/312/EXT https://wavenet.cefas.co.uk/details/312/EXT
- wslh 1y agoIn Google Chrome, at least, I tried an infinite loop modifying document.title and it freezes pages in other tabs as well. Now, I am not at my computer to try again.
- deleted 1y ago[deleted]
- jillyboel 1y ago[flagged]
- seventh12 1y agoThe intention is to crash bots' browsers, not users' browsers
- jillyboel 1y ago[flagged]
- anthk 1y agoIf you are crashing some browser from a disallowed directory in robots.txt, is not your fault.
- chrismorgan 1y ago[flagged]
- BeFlatXIII 1y ago> If you’re not familiar with this, read up on it, the reasons can be quite thought-provoking Are the reasons relevant to headless web browsers?
- chrismorgan 1y agoSome, definitely not. Others, quite possibly.
- bryanrasmussen 1y agohere's a potentially relevant example https://news.ycombinator.com/item?id=43947910 https://news.ycombinator.com/item?id=43947910
- 1y ago
- neuroelectron 1y agoI, for one, find it hilarious that "headless browsers" are even required. JavaScript interpreters serving webpages is just another amusing bit of serendipity. "Version-less HTML" hahaha
- kevin_thibedeau 1y agoIt exists because adtech providers and CDNs punish legitimate users who don't execute untrusted code on their property.
- Thorrez 1y agoHeadless browsers exist because adtech providers and CDNs punish legitimate users who don't execute untrusted code on their property? If we ask the creators of headless chrome or selenium why they created them, would they say "because adtech providers and CDNs punish legitimate users who don't execute untrusted code on their property"?
- wraptile 1y agoI find the "don't let googlebot see this" kinda funny considering how top google results are often much worse. The captcha/anti-bot is getting so bad I had to move to Kagi to block some domains specifically as browsing contemporary web is almost impossible at times. Why isn't google down ranking this experience?