4 ms·
Interesting: I posted this a few days back, and certainly not "10 hours ago". Who was kind enough to re-surface this? Thanks :) Some clarifications on a few co
by fforflo 2y ago
Interesting: I posted this a few days back, and certainly not "10 hours ago". Who was kind enough to re-surface this? Thanks :)
Some clarifications on a few comments I see downstream:
The motivating example was to easily support Full-Text Search (FTS) on PDFs with SQL only (see blog post https://tselai.com/full-text-search-pdf-postgres https://tselai.com/full-text-search-pdf-postgres ).
You can treat `pdf` as an alias for `text` and do everything possible.
On the next iteration, I made `pdf` a type (typical varlena object of bytes) to avoid hitting disk all the time. The file is loaded from the disk only once (if it's a valid pdf). One can store the `pdf` type (blob of bytes) as a standard Postgres type. And use that for subsequent calls. Postgres will do it's magic as usual.
There is a potential next step of storing the parsed document just to save some time from re-parsing the bytes, but I deemed it a premature optimization.
- scrlk 2y ago> Interesting: I posted this a few days back, and certainly not "10 hours ago". Who was kind enough to re-surface this? Thanks :) It's HN's second-chance pool: https://news.ycombinator.com/item?id=11662380 https://news.ycombinator.com/item?id=11662380