3 ms·
I’d be very interested in the opposite! Lots of scanned or legacy images that would be nice to convert to LaTeX, or to create a robust PDF ingestion pipeline.
by ComputerGuru 2y ago
I’d be very interested in the opposite! Lots of scanned or legacy images that would be nice to convert to LaTeX, or to create a robust PDF ingestion pipeline.
- Wdorf 2y agoI could add a GPT based pdf to latex functionality in the future.
- sorenjan 2y agoLike Mathpix? https://mathpix.com/ https://mathpix.com/
- BeetleB 2y agoMathpix is awesome. So far it has never gotten the output wrong. I even have it integrated into Emacs/org-mode.
- hextex 2y agoFacebook's Nougat [1] should work with this, but not sure how much preprocessing is needed to yield good results with scanned copies of physical documents. Note that it outputs .mmd files (MultiMarkDown), but the equations and tables should (iirc) output plain LaTeX. 1: https://github.com/facebookresearch/nougat https://github.com/facebookresearch/nougat
- Wdorf 2y agoThis looks really interesting! I will definitely have a look.
- Vetch 2y agoIn addition to the already mentioned https://huggingface.co/facebook/nougat-base https://huggingface.co/facebook/nougat-base, I also highly recommend https://huggingface.co/stepfun-ai/GOT-OCR2_0 https://huggingface.co/stepfun-ai/GOT-OCR2_0. It might even be better.