3 ms·
I recently redesigned my stack of validating uploaded files and creating thumbnails from them. My approach is to have different binaries per file type (currentl
by dividuum 2y ago
I recently redesigned my stack of validating uploaded files and creating thumbnails from them. My approach is to have different binaries per file type (currently images JPEG/PNG, videos H264/265 and truetype fonts). Each of them is implemented as in a way that they receive the raw data stream via stdin and then either generate an error message or a raw planar RGBA data stream via stdout. The validation and thumbnail process is triggered after first locking in the process in a seccomp strict mode jail before touching any of the untrustworthy data. Seccomp prevents them from basically every syscall except read/write. Even if there would be an exploit in the format parser, it would very likely not get anywhere as there’s literally nothing it could do except write to stdout. Outside a strict time limit is enforced.
The raw RGBA output is then received and converted back into PNG or similar. It was a bit tricky to get everything working without additional allocation and using syscalls triggered by glibc somewhere, but works pretty well now and is fast enough for my use case (around 20ms/item).
- artwr 2y agoOh could you expand briefly on what the stack looks like to accomplish this? Or do you have a write up on a blog/site you could share?
- dividuum 2y agoI wrote a little bit more about this here: https://community.info-beamer.com/t/an-updated-approach-to-content-sandboxing/1277 https://community.info-beamer.com/t/an-updated-approach-to-c... I’m a huge fan of building minimal self-contained tools, so all of the C programs statically link in the required parser libs (libavcodec/wuffs/freetype) so the resulting binaries don’t require additional dependencies on the target machine. The python wrapping code is rather straightforward as well and is only like 300 lines of code.