3 ms·
Out of curiosity, I tried submitting the first 200 pages of the PDF he used to my new tool that I also submitted today [0] to Show HN, ( fixmydocuments.com ), a
by eigenvalue 2y ago
Out of curiosity, I tried submitting the first 200 pages of the PDF he used to my new tool that I also submitted today [0] to Show HN, ( fixmydocuments.com ), and it generated the following without any further interaction besides submitting the PDF file:
https://fixmydocuments.com/api/hosted/m-moires-de-saint-simon-nouvelle-saint-simon-louis-02d09f https://fixmydocuments.com/api/hosted/m-moires-de-saint-simo...
I think it's not a bad result, and any minor imperfections could be revised easily in the markdown. My feature to turn the document into presentation slides got a bit confused because of the French language, so some slides ended up getting translated into English. But again, it wouldn't be hard to revise the slide contents using ChatGPT or Claude to make them all either French or English:
https://fixmydocuments.com/api/hosted/m-moires-de-saint-simon-nouvelle-saint-simon-louis-1eac5b https://fixmydocuments.com/api/hosted/m-moires-de-saint-simo...
[0] https://news.ycombinator.com/item?id=42453651 https://news.ycombinator.com/item?id=42453651
- bambax 2y agoThanks, but I'm sorry to say, the result is... bad. It invents words ("rusticas" in the second line of the title of the output, isn't anywhere in the source file -- it's not even a French word). And it completely drowns the footnotes inside the main text, inventing layout and text enrichment in the process. Footnotes are an important part of this project, if not the main point. If they are mangled with the main text then it's pointless. In your rendering there doesn't seem to be footnotes at all? just text with random titles here and there, and even more random tables that (to me) make no sense. I wouldn't call that "minor imperfections". As it is, it really isn't usable.
- eigenvalue 2y agoOK, thanks for the feedback. I really only tested with English language input documents, so I had low expectations going in. And you're right that this is certainly a challenging case with a lot of document structure and not very high quality scans.