4 ms·
A few things turn me off about Scrapy is that it feels over engineered for what it does. Why do I need an entire framework? I'm taking on technical debt to acc
by brilliantcode 10y ago
A few things turn me off about Scrapy is that it feels over engineered for what it does. Why do I need an entire framework?
I'm taking on technical debt to access data I don't have programmatic access to.
CSS/Xpath are very fragile. You most likely will be changing them in the future.
- eli 10y ago> CSS/Xpath are very fragile. You most likely will be changing them in the future. Genuinely curious what the alternative is
- brilliantcode 10y agoI've been doing research on this but it's not clear whether this problem is a pain for enough number of businesses to justify further investments. I often feel like web scraping is a commodity without understanding any of the inherent technological complexities and challenges. Very discouraging field to be in, especially when people claim to have pain but are unwilling to pay very much for it or show appreciation for the effort that goes into it. edit: thanks for the downvotes. perfect illustration of how innovation is punished and unrewarded in this field.
- jlgaddis 10y agoFYI, I only downvoted you after you complained about downvotes.
- brokenmachine 10y agoI only downvoted because there was no alternative offered, just complaining about how underappreciated scraper creators are. > perfect illustration of how innovation is punished I see no innovation in your post. The complaining about downvotes was just the icing on the cake, cementing the downvote.
- skinnymuch 10y agoHow many downvotes did you get? A few may have just been because your response is vague and doesn't say much for the question. But that doesn't mean a downvote should occur, of course. Otherwise complaining about downvotes is no good either. Some will downvote because of that.
- deadbunny 10y ago> especially when people claim to have pain but are unwilling to pay very much for it or show appreciation for the effort that goes into it. Welcome to pretty much every profession in the world.
- nikcub 10y agofor unstructured data applying NLP the other alternative is parsing schema.org schemas or other markup
- IanCal 10y agoI've used scrapy a few times and it's never felt like a big over-arching framework. I've been able to change what I need, and what it does for me are all the things I'd have to have done myself (caching, parallel requests but throttling per domain, dropping into debug mode, scheduled runs, etc). Really, it feels more like a skeleton + lots of sensible defaults. The meat of the code will be in the parsing, and so if for some reason I really needed to move away from it then I wouldn't feel particularly tied to it.
- brilliantcode 10y agoInteresting. So the web crawling/page fetching component is a major value add to you as a developer? Whereas the parsing is less of a value add because you prefer to code the parser yourself so that you have more control? What about changing the parser and crawler as the websites changes? What other pain points about scrapy do you have?
- IanCal 10y ago> Interesting. So the web crawling/page fetching component is a major value add to you as a developer? When I have a scraping task, yes. It's a set of things that are required each time but also a bit fiddly to get right, and scrapy has solved them well. > Whereas the parsing is less of a value add because you prefer to code the parser yourself so that you have more control? These I see as fundamentally having to be things I code as they're the parts that are different each time. > What about changing the parser and crawler as the websites changes? Really just a cost of doing scraping. Anything fancier has generally taken more time and created more problems than just assuming an ongoing maintenance cost. The debugging shell in scrapy is very useful for this. Last big job I did I also built a cache that you could query by time, so all versions of the page seen were stored which was very useful for debugging intermittent problems, and finding page changes. I don't know if scrapy has this in its cache, I don't think so but wouldn't conflict with it.