3 ms·
> If you find yourself in 2022 or later designing a file format intended for bulk data and you use any of the words "stream", "serialization", Nit: I think I k
by kortex 4y ago
> If you find yourself in 2022 or later designing a file format intended for bulk data and you use any of the words "stream", "serialization",
Nit: I think I know what you are getting at, with blocks streams vs byte streams, but it's kinda hard to design a file format without serialization or byte streams. Not sure how that would work.
> I think it's high time that the industry standardised on a generic "container" format to replace legacy archive file formats.
I have a side project chipping away at just such a thing. It's quite daunting, so if this at all interests anyone, please comment/reach out. I'd love more of an excuse to work on this.
SITO in a nutshell:
- It's all based on msgpack, which does most of the heavy lift for serialization and datatype encoding
- a sito stream comprises blocks, each block is a self-contained, independently decodeable msgpack array object.
- each block an array of the form (type: smallint, header: optional(hashmap), data: any)
- the type is either a single-byte int, or a packed int indicating a substream id
- the sito primary stream comprises multiple independent substreams
- there's no raw plaintext fields, but there is a plain unicode block which can be used to embed whatever plaintext metadata
- substreams can each have whatever compression/codec they want
- since each stream is a block stream, it's trivial to de/interleave
- for data integrity, I want to do something like block-level CRC/FEC along with per-stream merkle trees but I haven't worked out the details yet
- there are periodic "sync-blocks" which have a magic 8-byte sequence for starting a file, but also throughout the stream, to facilitate re-alignment of read heads
- I've also been toying with the idea of using sqlite as stream indexes and as a general glue to keep track of what's going on (right now, you can arbitrarily start a new substream at any point in the primary stream, so it's hard to tell at the start of a file what's in it, sqlite pre-allocates pages so write heads can go back and update a prior index block)
- nd-arrays are a particularly interesting datatype so there's an emphasis on ergonomics around handling them
I plan on doing a simple PoC at some point soon showcasing SITAR, the sito archive format, with a python tarfile-like interface.
- kortex 4y agoI just updated the github page with my current state of notes, in case folks are curious to some of the details. https://github.com/xkortex/sito https://github.com/xkortex/sito