6 ms·
Pdf.tocgen
- mrtx01 2y agoWhat a beautiful website!
- GrumpyNl 2y agoAnd build with very little CSS and basic HTML.
- oneeyedpigeon 2y ago> basic HTML Apart from the code blocks. Syntax-highlighting in `<code>` elements, when, browser manufacturers?
- porker 2y agoIt took a bit of digging from the Pdf.tocgen page, but https://krasjet.com/colophon/ https://krasjet.com/colophon/ tells us how it's created.
- lelandfe 2y agoUncommon to see someone so caring about the specifics of their chosen font. Love it.
- mbana 2y agoI love the typography on the site. What fonts are you using? I'm on a mobile browser so I can't really see.
- porker 2y agoAccording to https://krasjet.com/colophon/ https://krasjet.com/colophon/: > The typeface you are reading right now is Garibaldi by Henrique Beier, with some custom tweaks, as you might have noticed. I hope you enjoy it as much as I do. If you want some free alternatives, check out Alegreya ht and Vollkorn, though I still prefer the look and details of Garibaldi (just look at all the punctuation marks!).
- StayTrue 2y agoGaribaldi, $300 for up to 10k page views per month.
- karma_pharmer 2y agoI was going to post the same thing. This has to be the most beautifully typeset webpage I've seen in quite a while. Not just the font but the layout too. It's almost like this page is part of the web from some parallel universe, which has been disenshittified to the same extent that our own web has been... well, you know.
- papichulo2023 2y agoLooks like a very good tool to integrate with Knowledge Graphs or just RAG (llm).
- perihelions 2y ago- "That is, you shouldn’t expect it to work with scanned PDFs" It's surprisingly easy to extend this type of workflow to scanned pdfs (as opposed to software-generated, text-containing ones). tesseract(1) makes short work of ToC pages with --psm set to 6 (an OCR setting that tends to collapse convoluted text layouts into a regular, software-parseable output). It should also be straightforward, but I don't know of an out-of-the-box solution, to automate that example of extracting "text that looks like a header"–based on page layout/relative positioning, or font weight. (I'm working on an adjacent problem, an automatic re-layout of raster documents to squeeze out whitespace and make them slightly nicer on small e-ink devices. Text islands are trivial to identify. I don't know how to quantify font weight, or things like that. I'm "wasting" a lot of time diving into lots of mathematics rabbit holes, but I don't know in advance which ones will be productive or not).
- felipefar 2y agotesseract is fine for basic use cases, but it fails when the image is tilted (and thus the text isn't laid out horizontally), which can happen several times with scanned books. Compared to how well the Google OCR engine works, tesseract should be much better than it is. I wonder how difficult it is to develop a better OCR engine than tesseract.
- perihelions 2y agoAm I overlooking something, or is automating page rotation no more work than just a 2d FFT?
- notyoutube 2y agoMind ELI5ing this? it seems neat
- perihelions 2y agoThe Fourier transforms map plane waves to points. Blocks of regularly-spaced text have a periodic character, with the period length of their line spacing; their Fourier transform (I think??) would, in 2d frequency space, have amplitude peaks on vectors that have the same angle as the rotation of the lines.
- bionade24 2y agoDoes someone know a tool that is sed- or awk-like for PDFs?
- perihelions 2y agopdftk is a CLI tool that can extract and edit PDF metadata such as tables of contents*, if that's what you mean? *(Table of contents? Tables of content?)
- manaskarekar 2y agoPerhaps you can use lesspipe with sed/awk? https://github.com/wofr06/lesspipe https://github.com/wofr06/lesspipe
- maxerickson 2y agoQpdf has tools that go in that direction (but not a flat text format that allows arbitrary edits). https://qpdf.readthedocs.io/en/stable/qdf.html#qdf https://qpdf.readthedocs.io/en/stable/qdf.html#qdf
- janpmz 2y agoRecently I found the getToc function in PyMuPdf was too slow. I told them about it in their discord, and a day later they had fixed it. Now it only takes a couple of milliseconds. I'm using it for my project pdftomp3. Pdf.tocgen looks useful too, but I'm not sure if I can use it because of the licencse?
- zerop 2y agoInterested to know what is pdftomp3?
- janpmz 2y agoYou can upload a PDF and convert the chapters into MP3s (either original text or simplified text). But for PDFs without a table of contents, you can only convert single pages.
- karma_pharmer 2y agoOf course you can use it. What you can't do is deny others the same freedoms the license grants to you.
- cge 2y agoThere does appear to be some licensing awkwardness here. The license is nominally GPLv3, but it says it is based on AGPLv3 projects. It also appears to misidentify (it may have been correct at the time) PyMuPDF as GPLv3 when that appears to actually be AGPLv3. My assumption is that using this would require complying with AGPLv3? There's the additional oddity that a portion of the repository (the recipes directory) is licensed under CC-BY-NC-SA, and so the repository is not fully open source. This is particularly confusing, however, as the functional content of the recipes directory appears to be mostly records of direct observations of parameter choices in external documents and tools, and so doesn't seem like it would be copyrightable at all, at least in the US.
- zerop 2y agoCan I use this tool to get toc for arxiv papers ?
- T3RMINATED 2y ago[dead]
- jbecke 2y agoWe (macro.com) have something similar but without the recipe part in our pdf/word processor. It works pretty well on numbered headings but not so well on non-numbered. We’re thinking of porting over to LLMs at some point.
- maCDzP 2y agoThat is a beautiful website. I got lost in it and it created a sense of wonder. Nice.
- deleted 2y ago[deleted]
- pseingatl 2y agoSince when do you need the hyperref package to generate a table of contents under LaTeX (as the author claims)? \tableofcontents does the job.
- chazeon 2y agoI have been thinking about this, but for a while now, I have settled on using ChatGPT's GPT-4v's multimodal capability to generate a text file containing the titles and pages based on screenshots of the TOC. After that, I used a pikepdf-based Python script to bake the TOC into the PDF I had. The upside, compared to Krasjet's approach, is that this works not only for text-based PDFs but also for scanned PDFs, even old scanned journal papers. The downside is that, before baking the TOCs, you need to make adjustments to the PDF as sometimes the empty pages are not included. You also need to calculate the offset for the prologs, cover, etc. I have a script for this kind of adjustment, but there always is manual intervention involved.
- rajaravivarma_r 2y agoIs it possible to extract different patterns of text from a PDF document? For example, paragraphs, code blocks, code inlined in paragraphs etc? I tried tesseract but it recognises code blocks as tables. Also there are edge cases like paragraphs starting with an indentation and without an indentation are hard to differentiate. Appreciate any help.