19 ms·
Scraperr – A Self Hosted Webscraper
- iSloth 1y agoInteresting, wish it had markdown output like firecrawl for embedding/llm use cases
- smartmic 1y agoMy preferred "self-hosted" webscraper is a local, single binary called xidel [1]. The feature I really like is that it can also follow links. [1] https://github.com/benibela/xidel https://github.com/benibela/xidel
- darkwater 1y agoWow, it's written in Pascal! That surely brings me to memory lane.
- DocTomoe 1y agoWith Pascal being my first "adult" language, not used in 20 years ... it is surprising how readable that code is. Makes me wish for such simpler times.
- benibela 1y agothat fits, I wrote the first code for xidel almost 20 years ago and it still uses Pascal because I didn't plan to change it but just wanted to show people what I programmed 20 years ago
- darkwater 1y agoWhat environment are you using nowadays to develop in Pascal? Editor, possible plugins etc?
- benibela 1y agoFreePascal / Lazarus There was actually just a large discussion thread in the Lazarus forum wondering about why it is not more popular, with almost 300 comments. But now the thread got locked. That might be a reason. I used Delphi on Windows 98. But that became outdated, so I ported it to FreePascal. FreePascal has a lot of integrated libraries, but I do not really use anything I did not already use with Delphi.
- _QrE 1y agoIs there a reason for using Selenium over something like Playwright? I haven't had very many positive experiences with selenium, and playwright I found is easier to use and more flexible. Also, for stuff like this: `modified_value = original_value.replace("HeadlessChrome", "Chrome")` There's quite a few ways to figure out that a browser is a bot, and I don't think replacing a few values like this does much. Not asking you to reveal any tricks, just saying that if you're using something like Playwright, you can e.g. run scripts in the browser to adjust your fingerprint more easily.
- jpyles 1y agoI am quite aware, but I actually built most of the scraping logic a long time ago, before I even knew that playwright was a thing. I am looking to refactor a lot of this, and switching over to playwright is a high priority, using something like camoufox for scraping, instead of just chromium. Most of my work on this the past month has been simple additions that are nice to haves
- michaeljx 1y agoI was in a similar boat with my scrapers. Started with Selenium 5-6 years ago and only discovered Playwright 2 years ago. Spend a month or so swapping the two, which was well worth it. Cleaner API, async support.
- jpyles 1y agoWith the custom headers, you can actually trick a lot of sites with bot protection to let you load their sites (even big sites like youtube, which I have found success in)
- renegat0x0 1y agoNot a web scraper, but a web crawler software. Allows to specify method of crawling, selenium, and others. Returns data in JSON (status code, text contents, etc). [1] https://github.com/rumca-js/crawler-buddy https://github.com/rumca-js/crawler-buddy
- 3abiton 1y ago> extract data from websites with precision using XPath selectors. I've used XPath for crawling with selenium, and it used to be my favorite way, but turned out quite unreliable if you don't combine it with other selectors as certain website are really badly designed and have no good patterns. So what's the added value over pure selenium?
- cess11 1y agoCheck whether the site is actually server side rendered, because if it's a browser client that talks JSON to the backend, you could do the same.
- leelou2 1y ago[flagged]
- evertedsphere 1y agowould appreciate if you could post the prompt as well so the rest of us could learn how to generate our hn comments too
- nomilk 1y agoI used to scrape back in the day when it was easy (literally just make a request and parse html). Seems cloudflare checkboxes / human verification are very commonplace nowdays. Curious how(/if) web scrapers get around those?
- anxman 1y agoBy clicking the box
- welanes 1y ago1. Clicking the box programmatically – possible but inconsistent 2. Outsourcing the task to one of the many CAPTCHA-solving services (2Captcha etc) – better 3. Using a pool of reliable IP addresses so you don't encounter checkboxes or turnstiles – best I run a web scraping startup (https://simplescraper.io https://simplescraper.io) and this is usually the approach[0]. It has become more difficult, and I think a lot of the AI crawlers are peeing in the pool with aggressive scraping, which is making the web a little bit worse for everyone. [0] Worth mentioning that once you're "in" past the captcha, a smart scraper will try to use fetch to access more pages on the same domain so you only need to solve a fraction of possible captchas.
- nomilk 1y agoThat's awesome. Thanks for sharing. First time hearing of the fetch() approach! If I understand correctly, regular browser automation might typically involve making separate GET requests for each page. Whereas the fetch() strategy involves making a GET for the first page (just as with regular browser automation), then after satisfying cloudflare, rather than going on to the next GET request, use fetch(<url>) to retrieve the rest of the pages you're after. This approach is less noisy/impact on the server and therefore less likely to get noticed by bot detection. This is fascinating stuff. (I'd previously used very little javascript in scrapes, preferring ruby, R, or python but this may tilt my tooling preferences toward using more js)
- Tokumei-no-hito 1y ago
- lucb1e 1y agoFunny, I saw this HN headline just after banning another scraper's IP range You're welcome to scrape my sites but please do it ethically. Idk how to define that but some examples of things I consider not cool: - Scraping without a contact method, or at least some unique identifier (like your project's codename), in the user agent string. This is common practice, see e.g.: <https://en.wikipedia.org/wiki/User-Agent_header#Format_for_automated_agents_(bots) https://en.wikipedia.org/wiki/User-Agent_header#Format_for_a...>. Many sites mention in public API guidelines to include an email address so you can be contacted in case of problems. If you don't include this and you're causing trouble, all I can do is ban your IP address altogether (or entire ranges: if you hop between several IPs I'll have to assume you have access to the whole range). Nobody likes IP bans: you have to get a new IP, your provider has a burned IP address, the next customer runs into issues... don't be this person, include an identifier. - Timing out the request after a few seconds. Some pages on my site involve number crunching and take 20 seconds to load. I could add complexity to do this async instead, but, by having it live, the regular users get the latest info and they know to just wait a few seconds and everybody is happy. Even the scrapers can get the info, I'm fine computing those pages for you. But if you ask for me to do work and then walk away, that's just rude. It shows up in my logs as HTTP status 499 and I'll ban scrapers that I notice doing this regularly - Ignoring robots.txt. I have exactly 1 entry in there, and that's a caching proxy for another site that is struggling with load. If you ignore the robots file and just crawl the thing from A to Z at a high rate, that causes a lot of requests to the upstream site for updating stale caches. You can obviously expect a ban because it's again just a waste of resources
- edoceo 1y agoWhat do you have for log analytics and ban automation? Could you say more about how to identify these bad-bots?
- lucb1e 1y agoThere is no automation, I use `tail -f access.log` I just look at what's happening on my server every now and then. Sometimes not for months, but then when I set up a project like that caching proxy, I'm currently keeping a more regular eye to see that crawlers aren't bothering the upstream via me. Most respect the robots policy, most of the ones that don't set a user agent string that include the word 'bot' and so I know not to refresh the cache based on that request. So far it has mostly been Huawei who pretend to be a regular user but request millions of pages (from 12 separate IP ranges so far, some of them bigger than /16, some of them a handful of /24s). > Could you say more about how to identify these bad-bots? Many requests per day to random pages from either the same IP address (range), or ranges owned by the same corporation
- TheTaytay 1y agoDoes anyone know of a scraper that uses LLMs/natural language to build a deterministic, robust script that I can use to scrape the same site in the future? All of the natural language extractors I’ve seen so far need an LLM every time, but that seems unnecessary…
- throwup238 1y agollm-scraper [1] does a decent job but it's still a bit fragile. The biggest problem I have is all the React CSS-in-JS libraries that use hashes in their class names, which the LLM isn't smart enough to ignore. [1] https://github.com/mishushakov/llm-scraper https://github.com/mishushakov/llm-scraper
- TheTaytay 1y agoNice! Thanks!
- cdolan 1y agoWhat have you had success doing with this? Curious to test it
- throwup238 1y agoI mostly use it to aggregate event calendars for all the concert/sport/etc venues, meetups, and clubs in my area and do some other scraping tasks. I host a little wrapper around llm-scraper on a DigitalOcean droplet that I call from Val.town scripts I only check most places once a week so I use the LLM to do the scraping but there are a few cases where I have to scrape thousands of pages very frequently so I use the more deterministic script it generates instead.
- tengbretson 1y agoAnyone have any experience webscraping from a Starlink IP? My assumption is you could stay under the radar due to cg nat, but it's not exactly something I want to be the first to find out about.
- andrethegiant 1y agoShameless plug: prefix any URL with https://pure.md/ https://pure.md/ to get the pure markdown of that page. Useful for direct piping into an LLM. Has bot detection avoidance, proxy rotation, and headless JS rendering built in.
- matt-p 1y agoThat's excellent pricing from a structural perspective.
- fredoliveira 1y agothat looks fantastic - well done!
- yoble 1y agoLove the easter egg when going to https://pure.md/https://pure.md https://pure.md/https://pure.md
- vivzkestrel 1y agodoes this implement a rotating proxy IP address service?
- gitroom 1y agopretty cool seeing people still tweak their own scraping tools, but the cat and mouse game never ends huh - you think the web ever gets more open again or just keeps locking down?
- tommica 1y agoWell, it won't get more open by us just bitching here and doing nothing else
- jsemrau 1y agoI would prefer if we'd build a programmable web that provides value without relying on "scraping" websites for content. Most applications that do this are not well intended.
- gzkk 1y agoThere is quite high probability, that your own UserScripts will be well intended ;)
- monkeydust 1y agoPractical use-case. I am looking for a way to throw an address at a planning authority (UK) and download the associated documents for that property. Could this or another tool help? e.g. https://publicaccess.barnet.gov.uk/online-applications/applicationDetails.do?activeTab=documents&keyVal=SW1QTRJIIOQ00 https://publicaccess.barnet.gov.uk/online-applications/appli... As pure random example. A property can have multiple planning applications and under each many documents. What I have found useful (saved me time and potential lost £££) is to take the documents, combine to single pdf and provide to Gemini 2.5 Pro and then ask it to validate against agent specification for a property. Over the weekend found a place that was advertising a feature of the house that was explicitly prohibited through planning decision notice. Called the Agent up on it who claimed no knowledge but said this would have come up through solicitor checks, which it would have done, much later down the process with more or my money spent and considerable time lost. Of course all this possible without LLMs but just makes it easier/cheaper to check at scale.
- cess11 1y agoCould just cut out the href-value with grep and sed or a bit of scripting, '.pdf' seems to only occur on those links. I'd keep it simple like that until I need to do periodic comparisons, i.e. actually need scrapers and is prepared to build what's needed to automatically watch and process directories where the scrapers put the files.
- mellosouls 1y agoFrom the repo (clipped): When using Scraperr, please remember to: Respect robots.txt Terms of Service Rate Limiting Kudos for promoting ethical usage; makes a change from some of the grifters selling borderline ddos-bots as crawlers.