4 ms·
Do you plan on supporting scraping content via css selectors/xpath/regex?
by Toast_ 9y ago
Do you plan on supporting scraping content via css selectors/xpath/regex?
- RandomBookmarks 9y agoOr maybe integrating visual selection and OCR like Kantu: https://a9t9.com/kantu/scraping#ocr https://a9t9.com/kantu/scraping#ocr
- Toast_ 9y agoLooks cool, reminds me of portia. https://scrapinghub.com/portia/ https://scrapinghub.com/portia/
- onli 9y agoI'm not sure. There are a couple of existing services that do that already, and which give you a rss feed you could then import here. But if it gets requested I would look whether it is doable in the constraints of this system - at least I remember yahoo pipes supported it, that's a plus.
- Toast_ 9y ago>I remember yahoo pipes supported it, that's a plus. Yup, pretty much why I used pipes to begin with. I think being able to scrape/manipulate/output data, while being able to keep it private, would be a fantastic service. Looks good so far!
- onli 9y agoI added a block to download a page (instead of a feed), another block for extracting content via css selector or xpath, and a feed builder block to later combine those (but the extract block already creates a feed one could use as pipe output). If more is needed please say so, I am now convinced this fits well to this page.
- Toast_ 9y agoYo, good stuff on the added features. I'm currently using Huginn to scrape data, use a portion of that data to format a post request, and then combine the results of the post request with the scraped data, finally output as rss. Maybe some features to consider: ability to format get/post/put/delete request, ability to correctly (in order) merge objects (I haven't had the chance to try out your merge block yet). The merge, in my opinion, would be the biggest consideration as I have to use a custom agent to merge my Huginn events, and it's really a pain. Great start to the service man, keep up the good work! edit: I just tried the download agent on a site (http://www.plndr.com http://www.plndr.com) and it's throwing parse errors, and clicking the [x] won't close the output box, but the red portion works.
- onli 9y agoThanks :) I now understand the issue with the red portion and the [x]. That will be fixed soon. For the page, there was a bug with get params, those killed the output inspector. I fixed those now, it should be better able to fetch pages like http://www.plndr.com/product/browse?a=34714&catId=0&version=23296052e9cb4948fe055197a9b706e599815105&lg=1 http://www.plndr.com/product/browse?a=34714&catId=0&version=.... I was able to extract the product names from there in an example page, just a download block and an extract block selecting `.product-cell .product-title`. If you still have problems, would you please comment again, open a bug on https://github.com/pipes-digital/pipes/issues https://github.com/pipes-digital/pipes/issues or send me a mail? Kind of crucial to iron the kinks out. The parse errors are annoying, but I failed silencing them so far. The XML parser is throwing them regardless of try-catch, I don't know why. But they will be just ignored later on: 'View output' should show the pure html (instead of parsed and highlighted XML) instead. That seemed to work fine so far (but might fail in a different browser than those tested...).
- Toast_ 9y agoSure, I'd be happy to help. On viewing the html, I'm not sure what software stack you're using, but maybe check out the riko[1] library, which was recommended in a parent comment. [1] https://github.com/nerevu/riko https://github.com/nerevu/riko
- reubano 9y agoYou can check out my library riko [1, 2]. While it doesn't have a slick GUI, it does support most of the original yahoo pipes (including xpath [3] and regex [4]). [1] https://github.com/nerevu/riko https://github.com/nerevu/riko [2] https://www.youtube.com/watch?v=bpn2G3TAAYY https://www.youtube.com/watch?v=bpn2G3TAAYY [3] https://github.com/nerevu/riko/blob/master/riko/modules/xpathfetchpage.py https://github.com/nerevu/riko/blob/master/riko/modules/xpat... [4] https://github.com/nerevu/riko/blob/master/riko/modules/regex.py https://github.com/nerevu/riko/blob/master/riko/modules/rege...
- Toast_ 9y agoHey, just saw your comment. Looks like a pretty slick library; would something like this run in a jupyter/azure notebook?
- reubano 9y agoMost definitely [1]. [1] http://nbviewer.jupyter.org/github/reubano/riko-tutorial/blob/master/Tutorial.ipynb http://nbviewer.jupyter.org/github/reubano/riko-tutorial/blo...