5 ms·
Show HN: A Python Spider System with Web UI
- bowlofstew 12y agoThat is a nice tool....nice work!
- meowface 12y agoThis looks really nice. The API seems more user-friendly than scrapy's.
- Immortalin 12y agoAny plans for a gui based web scraper interface similar to portia?
- binux 12y agoCurrently, yes and no. pyspider is running original python code, something like portia is a code generator (Apologize if I'm wrong, I have not use it). So it can been made as another WebUI module. But for flexible, I have no idea how to make it right currently. So, We have a css selector helper, but no plan for a complete tool.
- prht 12y agoI am not trying to offend you, but I really don't understand when someone says "yes and no". I hear it more and more these days. Is this becoming a cliche? It can be "yes" or "no", not both together. "yes and no" is "no" for me.
- smoe 12y agoDon't know about other languages, but in german this phrase is pretty common when there is no clear yes or no answer. Like "yes to some extend but not completely"
- zbb 12y agoTake a look at source code. The package hirarchy is not pythonic (use "libs" as top package is not a good idea).
- binux 12y agoagree
- redacted 12y agoWhat is the recommended way? (Serious question, I have larger projects that I would someday like to refactor into proper packages)
- ngoldbaum 12y agoBecause the name "libs" is now installed into the global module namespace. It's better to use a less generic name.
- iamtew 12y agoThere is also these guides that provide plenty of information on how packages work and best practices: https://packaging.python.org/en/latest/distributing.html https://packaging.python.org/en/latest/distributing.html https://github.com/pypa/sampleproject https://github.com/pypa/sampleproject
- binux 12y agoI have "organize the code using a single top-level package".
- paulhauggis 12y agoWhy isn't it a good idea? I have plenty of projects setup this way and it works well. It looks pretty well organized to me.
- vertex-four 12y ago
- mrmondo 12y agoNice project! I do wish it supported a PostgreSQL backend rather than (or as well as I guess) MySQL.
- OedipusRex 12y agoCan someone explain what this is?
- bjblazkowicz 12y agoHow's the performance compared to scrapy?
- binux 12y agohttps://gist.github.com/binux/67b276c51e988f8e2c31 https://gist.github.com/binux/67b276c51e988f8e2c31
- adam-_- 12y agoHow does this compare to scrapy? Why would I use one over the other, or is either a fine choice?
- skillachie 12y ago+1 Scrapy comparison please Can you compare to scrapy as requested by other posters. Why could you not build on top of scrapy and leverage celery for scheduling etc (http://www.celeryproject.org/ http://www.celeryproject.org/) What is the immediate value add to using pyspider ?
- binux 12y agoI'm working on a benchmarking suite https://gist.github.com/binux/67b276c51e988f8e2c31 https://gist.github.com/binux/67b276c51e988f8e2c31 and meet some problem... pyspider comes from a vertical search engine project. we have two issues: - 100+ websites, they may change the template or down sometime. We need a dashboard to monitor the changes and the fails. - update in 5 minutes, when the website updated, we need follow that in 5 minutes. We are using a update time from index(list) page to tell the changed pages. And pages should been updated after about 30 days in case of we missed something. A powerful scheduler is needed. obviously, I hadn't got the right way to do so with scrapy. I'm not very familiar with scrapy. So I can't say something pyspider can do but scrapy not.
- kidsil 12y agoThanks for making me feel bad about my python-based aggregation solution :) https://github.com/AZdv/agricatch https://github.com/AZdv/agricatch
- erikb 12y agoWhat is a "spider system"? Never heard that term before.
- binux 12y agosorry :(
- deleted 12y ago[deleted]
- _bitliner 12y agoI really like the flow/UX. Congratulations! Nice job! What is the roadmap? I am really inside scraping, it is one of my daily job. I could consider to integrate it in one of my architectures
- _bitliner 12y agoFurthermore, what you mean with `Javascript pages supported`? Could I just specify where it has to click or do I need to make a reverse engineering of the ajax calls?
- binux 12y agohttp://demo.pyspider.org/debug/js_test_sciencedirect http://demo.pyspider.org/debug/js_test_sciencedirect is a sample for this. There is a phantomjs fetcher that can render the page as WebKit did. Furthermore, you can have some JavaScript running before/after page loaded to simulate a mouse click.
- binux 12y agoTo make it more flexible and easy to reuse? I have implemented most features I need now.
- _bitliner 12y agoBecause I already have a powerful distributed architecture. I was curious about the architecture of pyspider. For example, how the queue is handled? Is it centralized? Is there a server managing it?