Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
deanmalmgren
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
Scrubadub PII from dirty dirty unstructured text
(scrubadub.readthedocs.io)
12 points
by
deanmalmgren
9y ago
|
1 comments
2.
▲
by
deanmalmgren
12y ago
Hopefully it will be? There's a great suggestion to use tesseract-ocr to make this happen. https://github.com/deanmalmgren/textract/issues/16 If you have any other (better?) ways of doing this, feel free
3.
▲
by
deanmalmgren
12y ago
Thanks for the suggestion. I wasn't familiar with ps2ascii and I just created an issue here https://github.com/deanmalmgren/textract/issues/25
4.
▲
by
deanmalmgren
12y ago
Currently 2.7 but there's no reason python 3 can't be supported too. Thanks for the heads up on the borking of the pypi page. Noted.
5.
▲
by
deanmalmgren
12y ago
For what its worth, textract (python) also has ambitions of including OCR through the tesseract-ocr project https://github.com/deanmalmgren/textract/issues/16
6.
▲
by
deanmalmgren
12y ago
I thought about the metadata thing but decided to exclude it for the earliest versions of textract to keep things simple. If you'd like to see it in there and have a good example of how you'd like to use metadata, please feel free