3 ms·
Looks cool from what I’ve seen (well done on the release) of it (I read the read me and poked through your code). I’d be interested to see how this does agains
by makingstuffs 2y ago
Looks cool from what I’ve seen (well done on the release) of it (I read the read me and poked through your code).
I’d be interested to see how this does against sites which have things like Cloudflare’s bot detection enabled and sites such as Google Trends.
From my experience the stock version of playwright doesn’t play all too well with them and for sites like google trends there are a lot pain points when trying to liberate data.
Still interesting nonetheless
- techn00 2y agoIt doesn't, you cant find any reference to the captcha solving or anything like that in the repo. There's one line for the stealth plugin that is commented out: // chromium.use(stealthPlugin()); imo this is the hardest part about scraping, evading bot detection and captchas. edit: and keeping the scraping logic & rules up to date
- krsma 2y agoHi, creator here. So our open-source version does not provide cloudflare bypass or captcha support. It is impossible to have a robust system for the same completely FOSS. But we do have it available in our cloud version (which we launch soon, currently in testing). Our open-source version allows you to BYOP (Bring Your Own Proxy) to handle all the bypassing. Our OSS version is being used by users with small to medium scraping needs :) All of this is mentioned in our README.md.
- krsma 2y agoHi, creator here. So our open-source version does not provide cloudflare bypass or captcha support. It is impossible to have a robust system for the same completely FOSS. But we do have it available in our cloud version (which we launch soon, currently in testing). Our open-source version allows you to BYOP (Bring Your Own Proxy) to handle all the bypassing. Our OSS version is being used by users with small to medium scraping needs :)
- makingstuffs 2y agoThat makes perfect sense, I know from experience that things like Google Trends scraping is an absolute pain in the backside and requires a lot of jiggery-buggery in order to get around their bot detection — a proxy alone will do nothing to bypass them due to the mechanisms they’ve implemented.