3 ms·
I love scraping web and produce structured data from web pages. The only downside of using XPath or similar extracting approach is necessity of constant mainten
by bbayer 11y ago
I love scraping web and produce structured data from web pages. The only downside of using XPath or similar extracting approach is necessity of constant maintenance. If I have enough knowledge about machine learning, I would like to write a framework that analysis similar pages and finds structure of data without giving which parts of page should be extracted.
- stummjr 11y agoMaybe you should give a try at Portia (http://scrapinghub.com/portia/ http://scrapinghub.com/portia/). It does exactly what you mean. You may also be interested in this library: https://github.com/scrapy/scrapely https://github.com/scrapy/scrapely