5 ms·
Anyone else feel let down by all the existing container formats out there, and feel this itch to like, write their own? I know, I know, that's like the worst po
by kortex 5y ago
Anyone else feel let down by all the existing container formats out there, and feel this itch to like, write their own? I know, I know, that's like the worst possible idea, xkcd #927 "now there are 15 standards", but like, what if?
What I'm looking for is basically something that ticks all the boxes of LOC's 7 Sustainability Factors. Is able to store arbitrary binary data, along with metadata, in a single file. Easy to pack/unpack, particularly in Python. I'm pretty well covered for image/audio/video formats, but none of those lend themselves to n-dim arrays.
Ogg, matroska, and HDF5 are all container formats which can ostensibly do this, possibly even well. Ogg and matroska are supposed to be able to support arbitrary data steams and metadata, but actually finding tooling to do the operations I want to do has yet to bear fruit. HDF5 is insanely complex, and I have some concerns about performance and data corruption risk.
Tar is tempting but it does some really goofy things, like ending an archive with 1kB of zeros.
What actually looks most promising as an existing "format" is the RIFF family of chunked encoding. Dead simple, python stdlib even has a reader. Only real downside is 2*32 ~= 4 gigabyte size limit, though this can be overcome by using a "stream of chunks" model.
Alternatively, a stream based on Msgpack, with some slight constraints, is very appealing.
Sqlite3, also very tempting.
If I did write such a "format", it would definitely be built on the primitives of either RIFF, Msgpack, or Sqlite-as-application-file, so it'd be less of a "file format" and more of "existing format with some opinions about schema".
Protip for anyone also thinking about making their own file format: make sure the API/format supports some sort of arbitrary metadata model! Apache Feather and Parquet are pretty recent, but don't really support writing metadata, you kind of have to hack around it. IMHO, this is a huge oversight, especially considering the internal metadata is literally just JSON-encoded fields.
Edit: actually just reread the follow-up to "moving away from hdf5" and it mentions ASDF format. This looks really promising. YAML metadata header with binary blocks. Normally it's not a good idea to start with UTF8 and later go binary, but if you know when to stop reading text, it's not the worst. Also, it's a name collision with the asdf version manager (and also anything else named after rolling the left hand home keys)
https://cyrille.rossant.net/should-you-use-hdf5/ https://cyrille.rossant.net/should-you-use-hdf5/
https://towardsdatascience.com/saving-metadata-with-dataframes-71f51f558d8e https://towardsdatascience.com/saving-metadata-with-datafram...
https://en.m.wikipedia.org/wiki/Resource_Interchange_File_Format https://en.m.wikipedia.org/wiki/Resource_Interchange_File_Fo...
https://msgpack.org/index.html https://msgpack.org/index.html
https://www.sqlite.org/appfileformat.html https://www.sqlite.org/appfileformat.html
https://asdf-standard.readthedocs.io/en/latest/ https://asdf-standard.readthedocs.io/en/latest/
- Evidlo 5y agoAlso Feather: https://github.com/wesm/feather https://github.com/wesm/feather
- kortex 5y agoI mentioned that. The problem with feather at the moment is readily writing metadata. I currently have to use a hacky workaround to write the metadata.
- kortex 5y agoOkay, I've been playing around with ASDF a bit and I kinda really like it! It's quite performant - about 40% slower than reading/writing TIFFs, but supporting arbitrary metadata and array sizes. Supports in-band compression with zlib (ok), bzip2 (slow AF), and lz4 (quite fast, only about 50% slower than no compression, so in line with lz4 compressing the raw data). There's a couple of squicky things, for one, it looks like the extension model may allow RCE behaviors by deserializing arbitrary classes. It looks like it uses a URI/registration based system so by default, it is limited in what types it can try to load, but it's plugin-based, so a library you import could theoretically expose you. But that seems relatively low risk. But overall this format is really cool. Supports linking across multiple ASDF files, which I can see getting messy, but this is an incredibly powerful technique and they do it in a sane way.
- WorldMaker 5y agoZIP? - Has readers/writers/APIs in almost every language under the sun - It's data tree is a file tree, leaving it very simple from an abstraction standpoint - No inherent arbitrary metadata mechanism but plenty of ideas (in the wild, even) for storing arbitrary metadata in JSON or XML or YML "files" side-by-side data files - Used in many examples in the wild (DOCX, XSLX, ODT, JAR, etc)
- kortex 5y agoI've considered it. If I went that route, I'd probably use zstandard instead of PKWARE .zip (even though python stdlib has the zipfile interface built in). Zst is waaay faster, the API is more modern and actually architected from the ground up, supports parallel block processing, it's got rfc8878. It's quite nice. I'd have to still implement the object model on top of the block storage. With low compression levels, you can actually tune things to get faster write times than raw bytes, depending on CPU speed vs IO speed. So that's cool. http://facebook.github.io/zstd/ http://facebook.github.io/zstd/