3 ms·
yeah, detection of file types is a fun. I worked on several implementation of very precise type detection (for use in email & web proxies), and there are so man
by alexott 5y ago
yeah, detection of file types is a fun. I worked on several implementation of very precise type detection (for use in email & web proxies), and there are so many low-level details that affect that detection. The biggest fun is detection of type for container-based file formats, like, OLE-based office files (office 95-2003), Zip-based formats (MS Office > 2003, OpenOffice/LibreOffice, jar/war/ear, ...) where you need to actually look into the file to determined the type. And to add complexity, the implementation should be very fast if you run inside the web proxy where type is detected for each request (at least twice for request/response stage, or more if you process an archive).
Extraction of content is a separate fun ;-)
- nixpulvis 5y agohow does `file` stack up?
- alexott 5y agoFile was slower by at least 10 times slower (it was ~10 years ago). Second thing, for example it didn’t work with compound objects. Especially for things like word embedded in excel, etc. Similar for XML-based types, zip-based, … Plus rules aren’t very flexible - I’d used lisp-like language to describe detection patterns - that made it quite flexible in describing complex detection logic.