3 ms·
Implement full-text search maybe?
by raxi 5y ago
Implement full-text search maybe?
- andyxor 5y agothat would be cool, the only problem is size of data, about half a petabyte of pdfs uncompressed (77TB compressed). OCR and indexing at this scale would be pretty costly.
- raxi 5y agoComputing power might be donated by volunteers. If there are volunteer software engineers, there must be volunteers to run a container.
- fho 5y agoAlso papers do not follow any formatting standard, as a human that does not bother us too much, but OCR will have a hard time making sense of multiple columns with random text above and below, figure captions that are separated by the main text with only a little bit of vertical space, etc... There are some projects that try to automate that, but so far all I have used needed some manual intervention.