4 ms·
The main advantage (for now) is that the library has a single interface for both HTTP and headless browsers, and bundled auto scaling. You can write your crawle
by jancurn 2y ago
The main advantage (for now) is that the library has a single interface for both HTTP and headless browsers, and bundled auto scaling. You can write your crawlers using the same base abstraction, and the framework takes care of this heavy lifting. Developers of scrapers shouldn't need to reinvent the wheel, and just focus on building the "business" logic of their scrapers. Having said that, if you wrote your own crawling library, the motivation to use Crawlee might be lower, and that's fair enough.
Please note that this is the first release, and we'll keep adding many more features as we go, including anti-blocking, adaptive crawling, etc. To see where this might go, check https://github.com/apify/crawlee https://github.com/apify/crawlee
- robertlagrant 2y agoCan I ask - what is anti-blocking?
- fullspectrumdev 2y agoUsually refers to “evading bot detection”. Detecting when blocked and switching proxy/“browser fingerprint”.
- robertlagrant 2y agoIs this a good feature to include? Shouldn't we respect the host's settings on this?
- nlh 2y agoIt’s a fair and totally reasonable question but clashes with reality. Many hosts have data that others want/like to scrape (eBay, Amazon, Google, airlines, etc.) and they setup anti-scraping mechanisms to try and prevent scraping. Whether or not to respect those desires is a bigger question but not one for the scraping library - it’s one for those doing the scraping and their lawyers. The fact is - many many people want to scrape these sites and there is massive demand for tools to help them do that, so if APIFY/Crawlee decide to take the moral ground and not offer a way around bot detection, someone else will.
- thebytefairy 2y agoAh yes, the old 'if I don't build the bombs for them, someone else will'. I don't think this is taking the moral high ground, this is saying we don't care whether it's moral, there's demand and we'll build it.
- amarcheschi 2y agoI'm not gonna feel bad if a corporation gets its data scraped (whenever it's legal to do so, and this is another kind of question I'm not knowledgeable enough to face) when they themselves try to scrape other companies' data
- robertlagrant 2y agoYou seem to have a massive category error here. To my understanding, this is not only going to circumvent the scraping protection of companies that scrape other people's data.
- jancurn 2y agoThere are many legitimate and legal use cases where one might want to circumvent blocking of bots. We believe that everyone has the moral right to access and fairly use non-personal publicly available data on the web the way they want, not just the way the publishers want them to. This is the core founding principle of the open web, which allowed the web to become what it is today. BTW we continuously update this exhaustive post covering all legal aspects of web scraping: https://blog.apify.com/is-web-scraping-legal/ https://blog.apify.com/is-web-scraping-legal/
- beeboobaa3 2y agoThoughts on this law? https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX%3A31996L0009 https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX%3A...
- mnmkng 2y ago