5 ms·
I remember back when Anubis came out, some naysayers on here were saying it wouldn't work for long because the scrapers would adapt. Turns out careless, unethic
by QuiDortDine 8mo ago
I remember back when Anubis came out, some naysayers on here were saying it wouldn't work for long because the scrapers would adapt. Turns out careless, unethical vibecoders aren't very competent.
- wolfi1 8mo ago"Turns out careless, unethical vibecoders aren't very competent." well, they rely on AI, don't they? and AI is trained with already existing bad code, so why should the outcome be different?
- tuhgdetzhh 8mo agoI still think it is just a matter of time until scrapers catch up. There are more and more scrapers that spin up an full blown chromium.
- kstrauser 8mo agoIt seems inevitable, but in the mean time, that's vastly more expensive than running curl in a loop. In fact, it may be expensive enough that it cuts bot traffic down to a level I no longer care about defending against. Like GoogleBot had been crawling my stuff for years without breaking the site. If every bot were like that, I wouldn't care.
- raw_anon_1111 8mo agoSerious question, in 2026 you can actually have a successful crawler with just curl? I just had to create one for a customer - for their own site - and nothing would have worked without using Chromium.
- kstrauser 8mo agoProbably not for most sites. Example of a site where it'd likely work: a blog made with a static site generator. Example of one where it wouldn't: darn near anything made with React.
- franga2000 8mo agoIt works for the majority of things a text mining scraper would care to scrape. It's not just static sites but also any CMS like wordpress, as well as many JS apps that have server-side rendering. SPA-only sites aren't that common anymore, especially for things like blogs, news and text-based social media.
- hxtk 8mo agoEven that functions as a sort of proof of work, requiring a commitment of compute resources that is table stakes for individual users but multiplies the cost of making millions of requests.
- cantalopes 8mo agoWell it's a race, just like security. And as long as anubis is in the front, all looks bright
- gruez 8mo agoAFAIK you can bypass it with curl because there's an explicit whitelist for it, no need for a headful browser.
- solid_fuel 8mo agoCool, if they're running full blown chromium maybe the next step can be mining bitcoin on any pages served to bots.
- Elfener 8mo ago> Turns out careless, unethical vibecoders aren't very competent. Well they are scraping web pages from a git forge, where they could just, you know, clone the repo(s) instead.