3 ms·
Prompted by downloading a .doc file from Qwest only to find out that inside was a monospaced text file, I set up a small, nearly UI-free site for doing document
by afiler 15y ago
Prompted by downloading a .doc file from Qwest only to find out that inside was a monospaced text file, I set up a small, nearly UI-free site for doing document conversions. http://doc.mar.cx/<url> http://doc.mar.cx/<url> gives an HTML or other sensible rendering of an url (e.g. http://doc.mar.cx/http://www.itu.int/dms_pub/itu-t/oth/02/02/T02020000010001MSWE.doc http://doc.mar.cx/http://www.itu.int/dms_pub/itu-t/oth/02/02... ) and http://doc.mar.cx/<extension>/<url> http://doc.mar.cx/<extension>/<url> attempts to convert the url into the format with the given extension (e.g. http://doc.mar.cx/txt/http://www.itu.int/dms_pub/itu-t/oth/02/02/T02020000010001MSWE.doc http://doc.mar.cx/txt/http://www.itu.int/dms_pub/itu-t/oth/0... ).
I use wvHtml for doc->html, wvPDF for doc->pdf, but antiword for doc->txt. To convert .docx, .xls, .xlsx, and WordPerfect files to HTML, I use OpenOffice, by way of jodconverter. For ODF files, I use OdfConverter. Conversion of Excel files to .csv files uses xls2csv. For PowerPoint files, I use ppthtml to convert to html, and catppt to convert to text. For Lotus 1-2-3 files (I added this after downloading some historical telecom data from the FCC!), I use ssconvert.
Any conversion that results in an HTML file (e.g. doc or pdf to html) I bundle all the images into a single file using the data: url scheme. To do this, I wrote a utility called pagecan: http://afiler.com/pagecan/ http://afiler.com/pagecan/