3 ms·
Great work! I am a bit confused with the comparison with nougat throughout the repo. Nougat was specifically trained for academic documents, and I don't think a
by alsodumb 3y ago
Great work! I am a bit confused with the comparison with nougat throughout the repo. Nougat was specifically trained for academic documents, and I don't think anyone ever claimed Nougat was the best OCR model out there. That's kinda clear in your benchmark too where you mention that nougat has higher accuracy on arxiv documents. You also mention that marker will convert fewer equations when compared to nougat, and yet compare with nougat in terms of speed? (again, only complaining because it's a model designed for academic documents).
For anyone trying to do OCR on any pdf with math in it, definitely do try nougat. It's very easy to install (just a python package), and extracts the math, text, tables and beyond (in a .mmd file) with a single command line command. It also runs reasonably fast for personal uses - it takes about 30 seconds to convert a 6 page document using CPU only on my 4 year old i5 laptop.
- defsectec 3y agoHow do you think nougat would handle RPG rulebook PDFs? I'm looking for a food OCR model to help me transcribe sections of RPG books to markdown. Ideally, I'd like the emphasis such as bold or italics to be transcribed. The combo of text, numbers, and math symbols seems similar to technical and academic writing, but often has weird formatting, text boxes in the margins, and many diagrams.
- alsodumb 3y agoI'm not completely sure to be honest, but you should try it yourself with a sample page! I believe hugging face hosts it online on their demo pages so you don't even have to install the package to test on one page.
- fshr 3y ago> I don't think anyone ever claimed Nougat was the best OCR model out there Comparing two things doesn't inherently imply the previous thing was touted about with superlatives. It's just a way to juxtapose the new thing with something that may be familiar. As you said, nougat is easy to install/run so it makes sense they'd compare it. Would it be better if they could add more libraries in the comparison? Absolutely; that'd be helpful.
- vikp 3y agoAuthor here: for my use case (converting scientific PDFs in bulk), nougat was the best solution, so I compared to it as the default. I also compare to naive text extraction further down. Nougat is a great model, and converts a lot of PDFs very well. I just wanted something faster, and more generalizable.
- civilitty 3y agoGreat work! I just tried it on Linux for System Administrators and it did a great job properly picking up on code and config text. I noticed marker downloaded a PyTorch checkpoint called `nougat-0.1.0-small`, do you use nougat under the hood too or is that just a coincidence?
- vikp 3y agoYes, nougat is used as part of the pipeline to convert the equations (basically marker detects the equations then passes those regions to nougat). It's a great model for this.
- Ldorigo 3y agoReading your comment and parent's I think perhaps there is a mistake in the comparison chart on GitHub? It says nougat takes around 700 seconds per page and yours around 90. This doesn't match with parent's claim that it took him 30 seconds to run nougat on 6 pages.
- deleted 3y ago[deleted]
- sumedh 3y ago> extracts the math, text, tables I want to extract financial statements from pdfs which are in tables, would Nougat be suitable for that use case?