3 ms·
I would think you should reasonably be able to open those files with a regular text editor (vim comes to mind) and manually extract the contents .. right? I gu
by jesse__ 1y ago
I would think you should reasonably be able to open those files with a regular text editor (vim comes to mind) and manually extract the contents .. right? I guess if there was disk corruption and that produced an invalid UTF8 stream then maybe not .. but that'd at least be a smoking gun pointing to corruption, versus nobody being able to read the files anymore..
- mananaysiempre 1y agoIf you use a non-Latin alphabet, Microsoft Word’s RTF output is a horrific mess of encoding switches everywhere that makes manual text extraction pretty much untenable (and while RTF can use both UCS-2 and Windows codepages, Word seems to stick to—potentially multiple—codepages if it can, presumably for compatibility). That said, Microsoft always intended RTF to be Word’s exchange and archival format (unlike DOC, which was a mess they did not want to document), so it has enough of an official spec that extracting text, at least, is very possible.
- torstenvl 1y agoRTF uses UTF-16, not UCS-2; you can in fact use two \u____ commands in a row using surrogate pairs. Anyway, I wonder if this would work for you. https://github.com/torstenvl/rtfproc https://github.com/torstenvl/rtfproc