3 ms·
Show HN: Paradict – Streamable multi-format serialization with schema
Hi HN ! I'm Alex, a tech enthusiast. I'm excited to show you Paradict (https://github.com/pyrustic/paradict https://github.com/pyrustic/paradict), my solution for streamable multi-format serialization.
Although JSON, YAML, and TOML are all human-readable, they serve different purposes. For example, TOML is specifically designed for configuration files while JSON is used as a data interchange format.
Sometimes an initiative to create a binary version of JSON arises and as far as I know, it ends with an unidirectional mapping of datatypes.
There is no silver bullet, yet one coherent solution built from scratch that addresses multi-format (binary and textual) serialization and configuration files would be a step forward.
Earlier this year, I accidentally designed a textual data format to represent complex data structures inside a document divided into sections. The project, namely Jesth (Just Extract Sections Then Hack'em), generated an interesting discussion on HN (https://news.ycombinator.com/item?id=35991018 https://news.ycombinator.com/item?id=35991018).
Out of curiosity, I ran some benchmarks using Jesth, JSON and MessagePack, with and without Gzip compression against a large JSON file downloaded from the web. The benchmarking gave me insights that led to the decision to evolve Jesth's ideas into a new multi-format serialization solution.
I designed and built Paradict from scratch to serialize and deserialize a dictionary data structure. Although Paradict's root data structure is a dictionary, lists, sets, and dictionaries can be nested within it at arbitrary depth.
A Paradict dictionary can be populated with strings, binary data, integers, floats, complex numbers, booleans, dates, times, datetimes, comments, extension objects, and grids (matrices). There is also a schema-based validation mechanism that can contain programmatic checkers.
The binary serialization format is designed with compactness in mind such as Pi with its first two decimal places, the Golden ratio with its first two decimal places, and the date of the funeral of Pope Benedict XVI would each be encoded on two bytes (not counting their respective 1-byte tag which starts each Paradict binary datum).
This binary format has two levels of granularity for continuous data stream processing: a datum at the low level, which is in some cases a 2-tuple composed of a tag and its payload, and the message at the high level which is a dictionary data structure.
The textual serialization format has two modes: data and config modes. Config mode implicitly treats dictionary keys as strings, removing the need to surround them with quotes, and unlike the colon (:) between a key-value pair in data mode, it uses the equal sign (=) as separator.
This textual format has two levels of granularity for continuous data stream processing: a single line of text at the low level and the message at the high level which is a dictionary data structure.
Here is a valid Paradict configuration document that contains a "user" section:
[user]
# no comment
id = 42
name = 'alex'
birthday = 2042-12-25T16:20:59Z
photo = (bin)
54 68 69 73 20 69 73 20 6E 6F 74 20 61 20 70 68
6F 74 6F 67 72 61 70 68
weight_matrix = (grid)
1 0 1 0
0 1 0 1
1 0 1 0
books = (dict)
romance = (list)
'Happy Place'
'Romantic Comedy'
sci_fi = (list)
'Dune'
'Neuromancer'
epitaph = (text)
According to the law of conservation of energy,
no a bit of you is gone;
you are just less orderly.
---
Under the hood, Paradict uses Braq (https://github.com/pyrustic/braq https://github.com/pyrustic/braq), the most obvious way to section a document (as shown just above), and Ustrid (https://github.com/pyrustic/ustrid https://github.com/pyrustic/ustrid), to uniquely generate string identifiers.
Paradict is available on PyPI and you can learn more by reading its README, browsing the source code or playing with its tests.
Let me know what you think about all this !
- eviks 3y agoInteresting approach to various options for various use cases, though given such richness, for configs it's surprisingly poor with only a-z_ keys, and also doesn't seem to go far enough - great that you don't need quotes for keys, but why do you need quotes for values, is there no way to simplify data types a bit to be able to get rid of those?
- alexrustic 3y agoThank you for your comment ! > great that you don't need quotes for keys, but why do you need quotes for values ... Quotes for string values are very important because they avoid ambiguity [1][2] and the type of quotes (single or double) tells the deserializer how to treat the string: as an ordinary string (which may have escape sequences) or a raw string. > ... is there no way to simplify data types a bit to be able to get rid of those? String values already benefit from a simplification: quotes can be placed inside a string (ordinary or raw) without a backslash to escape them. This is not the case in TOML: "Since there is no escaping, there is no way to write a single quote inside a literal string enclosed by single quotes. Luckily, TOML supports a multi-line version of literal strings that solves this problem." (https://toml.io/en/v1.0.0#string https://toml.io/en/v1.0.0#string) > for configs it's surprisingly poor with only a-z_ keys Config dictionary keys follow the identifier naming rules in C to enable mapping between config keys and variables. Therefore, one could (at least in Python) pass a config dict as arguments to functions like this: from paradict import ConfigFile def my_func(arg1=42, arg2=True, **kwargs): pass # load user_config from the 'config.dict' file path = "/path/to/config.dict" confile = ConfigFile(path) user_config = confile.get("user") # pass user_config to my_func my_func(**user_config) [1] https://news.ycombinator.com/item?id=30052128 https://news.ycombinator.com/item?id=30052128 [2] https://news.ycombinator.com/item?id=28826600 https://news.ycombinator.com/item?id=28826600
- eviks 3y agoIn a huge number of cases, especially end user app config ones, not those production Norway ones, this ambiguity doesn't matter, and could be resolved through user choice of always quoting despite it being non-mandatory Re raw string - so let only it be quoted (or maybe better yet - vice versa, no quotes, no expansion), that's a valid mechanism to signal type. Just like you'd need a quote if you wanted to add whitespace are the beginning of a regular string (unless you make it only 1 space after = a valid separator) I don't get the a-z benefit in the argument case - the user must type "arg1" precisely for the argument names to match, so how does allowing Unicode make it different?