4 ms·
I scan all documents and paper mail and store the scans in a git annex repo with a few replicas on machines I own or rent. Originals go into binders, ordered by
by schube 3y ago
I scan all documents and paper mail and store the scans in a git annex repo with a few replicas on machines I own or rent. Originals go into binders, ordered by the hash of the document. With a few helper scripts and OCR with Tesseract searching through my documents is surprisingly fast, so is adding new documents at the proper place in the binders. The idea behind ordering by sha1 of the scan was that the binders would be storage only and all categorizing and organizing should be digital, and also that as I add more documents load balancing between the binders is trivial: at the moment I have one binder for hash prefixes 0-7 and one for 8-f, with a divider for each one-hex-digit prefix. If I ever feel there are too many docs per divider I can split further.
Emails go into a plain git repo, and photos go into another git annex repo.
Ideally I’d like to put all of that into Perkeep because I like the idea of just throwing anonymous blobs into a bucket with the possibility to organize and tag them after the fact, but it looks like that project is dead and I would need to add a few things to it to support my use cases.
I store a lot of scans and emails I won’t ever need again. I just consider that storage becomes cheaper faster than my storage needs increase, and that storing extra documents and saving myself the need to decide whether it’s worth keeping, is a good tradeoff.