10 ms·
Show HN: Sukhoi – A flexible and extensible Webcrawler in Python
- monksy 9y agoHow does this differ from Scrapy?
- iogf 9y agoThe way of how you construct your json structures in scrapy it is different, scrapy has a longer learning curve too. It seems sukhoi has got better results in performance too.
- iogf 9y agoTry to imagine how to solve the second example of the sukhoi README.md using scrapy, you'll notice you'll end up with some kind of obscure logic to achieve that json structure thats outputed by the second example in sukhoi's README.md.
- bbernoulli 9y agoFWIW I don't believe this would be overly convoluted in scrapy. I'd probably scrape the tags and quotes in one pass... Also, generator expressions would make the examples more readable IMO. self.extend((tag, QuoteMiner(self.geturl(href))) for tag, href in self.acc)
- iogf 9y agoI would like to see that in scrapy. I think you may have a point about the generators, yea.
- vosper 9y agoHow useful are scrapers that don't execute Javascript these days? I find Selenium + PhantomJS (now Chrome Headless I guess) is pretty easy to drive from Python, and it works everywhere because it's a real browser.
- dhruvkar 9y agoPhantom still gets blocked, as it reveals itself in the header. I've had success with a headless Chrome instance in a virtual display (xvfb) driven with Selenium, backed by Postgres. It's as close you can get to scripting a real browser.
- rectangletangle 9y agoYou can set the user-agent with Phantom var webPage = require('webpage'); var page = webPage.create(); page.settings.userAgent = 'Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/37.0.2062.120 Safari/537.36'; http://phantomjs.org/api/webpage/property/settings.html http://phantomjs.org/api/webpage/property/settings.html
- dhruvkar 9y agoWhile that's true, user-agent isn't the only thing in the header that reveals PhantomJS[0]. You could take the time to build in spoofs for these issues. But for testing (and scraping), you're going to be better off if your headless browser is the same as your GUI browser. 0: https://blog.shapesecurity.com/2015/01/22/detecting-phantomjs-based-visitors/ https://blog.shapesecurity.com/2015/01/22/detecting-phantomj...
- eriknstr 9y agoI think I read recently that Chrome / Chromium is now able to run without having to use xvfb, so now truly headless.
- iogf 9y agoThese ones are slower in most cases, it seems for some situations the ones that dont execute js would better do the job.
- rectangletangle 9y agoThey're still surprisingly useful, however it depends a lot on your use case. In my case I've scraped quite a bit from Wikipedia (not everything is available in a clean API) and other sources this way.
- dguo 9y agoInteresting timing! I just started using Scrapy today for a project, and I'm trying to figure out how to elegantly piece together information from different sources. I'm glad to see that that problem is the focus of your README example.
- tedmiston 9y agoPretty cool project. It looks more enjoyable to use than BeautifulSoup. How does it approach throttling or rate limiting? I didn't see this mentioned in the readme examples. Would be nice if there were some simple config to kick requests back into a queue to be re-run once limits aren't exhausted. Minimal support for caching / ETag / etc would be a nice addition.
- iogf 9y agoThe throtting can be set directly from untwisted reactor(planning to implement soon once i get untwisted on py3). I think the support for caching is really good too, i plan to implement it this week.
- tedmiston 9y agoAwesome. It looks like you're reusing your own dependencies which is cool. Can you explain how untwisted relates to twisted a little more? I read the repo readme, but not sure I'm following.
- iogf 9y agoUntwised is meant to solve all problems twisted solves but it does it in quite a different way. They are two different tools that would solve the same problems using different approaches. Untwisted doesnt share code nor architecture with twisted. In untwisted, sockets are abstracted as event machines, they are sort of "super sockets" that can dispatch events. You map handles to Spin instances, these handles are mapped upon events, when these events occurs then your handles get called. The handles can spawn events inside the Spin instances, in this way you can better abstract all kind of internet protocols consequently achieving a better level of modularity and extensibility. That is one of the reasons that sukhoi's code is sort of short, it is due to the underlying framework in which it was written on.
- sandGorgon 9y agoI'm not able to figure out dependencies.. is this pure python ? Or are you using one of gevent, libev, uvloop, etc. Since it is py2, i suppose asyncio is out of the picture
- pryelluw 9y agoIs this Python 3 compatible? Searched but the wiki is empty and the readme has examples in Python 2.
- deleted 9y ago[deleted]
- gear54rus 9y agoName seems to reference a prominent Russian aerospace engineer or maybe that's just wishful thinking. https://en.wikipedia.org/wiki/Pavel_Sukhoi https://en.wikipedia.org/wiki/Pavel_Sukhoi
- doubleplusgood 9y agoAlso it literally means "dry".
- ldng 9y agoDepending on your needs, sometimes it might be more interesting starting for there : https://about.commonsearch.org/ https://about.commonsearch.org/ and then scrap whatever is missing or not fresh enough. The scrapping process can be quite intense on servers.