3 ms·
The example on the Tesseract.js page shows it highlighting the rectangles of where the selected text originated. Does this level of information get surfaced thr
by fbdab103 3y ago
The example on the Tesseract.js page shows it highlighting the rectangles of where the selected text originated. Does this level of information get surfaced through the library for consumption?
I just grabbed a two-column academic PDF, which performed as well as you would expect. If I was returned a json list of text + coordinates, I could do some dirty munging (eg footer is anything below this y index, column 1 is between these x ranges, column 2 is between these other x ranges) to self-assemble it a bit better.
- simonw 3y agoYes it does, but I've not dug into the more sophisticated parts of the API at all yet. I'm using it in the most basic way possible right now: const {data: {text}} = await worker.recognize(imageUrl);
- deleted 3y ago[deleted]