6 ms·
Not open-source but free: Kantu (https://kantu.io https://kantu.io) uses OCR to support web scraping. You mark an anchor image/text with a green frame and mark
by RandomBookmarks 10y ago
Not open-source but free: Kantu (https://kantu.io https://kantu.io) uses OCR to support web scraping. You mark an anchor image/text with a green frame and mark the area of data that needs to be extracted with pink frames. The image inside the pink frames is then sent to https://ocr.space https://ocr.space for processing and Kantu api returns the extracted text. This works very well as long as you do not need a lot of data. It is certainly not a "high-speed" solution for scraping terabytes of data.
- brilliantcode 10y agoTried the OCR for scraping and gave up because it was too slow and inaccurate. OCR works well for certain scenarios where UI is fixed like on desktop applications but it's still fragile very much like CSS and Xpath selectors. In fact, often OCR performs far slower and less accurate than CSS/Xpath selectors. It has it's niches but I think it's sub optimal for web automation/scraping.