3 ms·
- this encodes to ASCII text (unless your strings contain unicode themselves) - that means you can copy-paste it (good luck doing that with compressed JSON or C
by creationix 7mo ago
- this encodes to ASCII text (unless your strings contain unicode themselves)
- that means you can copy-paste it (good luck doing that with compressed JSON or CBOR or SQLite
- there is a scale where JSON isn't human readable anymore. I've seen files that are 100+MB of minified JSON all on a single very long line. No human is reading that without using some tooling.
- bawolff 7mo agoThat kind of feels a bit worst of both worlds. None of the space savings/efficiency of binary but also no human readability. Being able to copy/paste a serialization format is not really a feature i think i would care about.
- creationix 7mo agoIt's a gradient. I did design several binary formats first, but for my use cases, this is actually better. There is nuance to various use cases. > None of the space savings/efficiency of binary For string heavy datasets, it's nearly the same encoding size as binary. I get 18x smaller sizes compared to JSON for my production datasets. This was originally designed as a binary format years ago (https://github.com/creationix/nibs https://github.com/creationix/nibs) and then later after several iterations, converted to text. > Being able to copy/paste a serialization format is not really a feature i think i would care about Imagine being paged at 3am because some cache in some remote server got poisoned with a bad value (unrelated to the format itself). You load the value in dashboard, but it's encoded as CBOR or some binary format and so you have to download it in a binary safe way, upload that binary file to some tooling or install a cbor reader to your CLI. But then you realize that you don't have exec access to the k8s pods for security reasons, but do have access to a web-based terminal. Again, to extract a binary value you would need to create a shell, hexdump the file and somehow copy-paste that huge hexdump from the web-based terminal to your local machine, un-hex dump it, and finally load it into some CBOR reader. A text format, however is as simple as copy-paste the value from the dashboard and paste into some online tool like https://rx.run/ https://rx.run/ to view the contents.
- rendaw 7mo agoAre there any examples? If it's ASCII I'd expect to see some of the actual data in the readme, not just API. Unless, to read that correctly, it only has a text encoding as long as you can guarantee you don't have any unicode?
- creationix 7mo agooh, sorry about that. I forgot to include the description of the format with examples. I did add some small examples to the repo. https://github.com/creationix/rx/blob/main/samples/quest-log.rx https://github.com/creationix/rx/blob/main/samples/quest-log... The older, slightly outdated, design spec is in the older rex repo (this format was spun out of the rex project when I realized it's actually a good standalone format) https://github.com/creationix/rex/blob/main/rexc-bytecode.md https://github.com/creationix/rex/blob/main/rexc-bytecode.md
- SV_BubbleTime 7mo ago'fdiscovered,aextreme,7danger,6+1A+16;6level_range,b:QThe Heap ,d'th Oof.
- dontdoxxme 7mo agoVery similar to bittorrent’s bencode. That has the benefit that it has a canonical encoding which this doesn’t (because of the different compression options). I wouldn’t be put off by how it looks as text.
- creationix 7mo agoVery true. I had forgotten about bencode, I should read up on that again. It makes sense they need a canonical form because they want same values to have same content hashes.
- creationix 7mo ago> it only has a text encoding as long as you can guarantee you don't have any unicode? The format is technically a binary format in that length prefixes are counts of bytes. But in practice it is a textual format since you can almost always copy-paste RX values from logs to chat messages to web forms without breaking it. unciode doesn't break anything since strings are encoded as raw unicode with utf-8 byte length prefixes. It supports unicode perfectly. If your data only contains 7-bit ASCII strings, the entire encoding is ASCII. If your data contains unicode, RX won't escape it, so the final encoding will contain unicode as UTF-8.
- kukkamario 7mo agoYou don't want to copy-paste anything like that as text anyway. Just copy and paste files. No human is reading much data regardless of the format. What is the benefit over using for example BSON?
- creationix 7mo ago> Just copy and paste files If all your workflows allow copying as binary files, more power to you! But there are a lot of workflows where that is not possible. This was inspired by years of hands-on operational incident handling in production systems. Every time we use a binary format, it's extra painful. This particular format would be slightly more compact as binary, but not enough to justify closing the door on all the use cases that would preclude. I'll probably add a binary variant for people who prefer that (or for people who want to be able to embed binary values in the data without base64 encoding it)
- mpeg 7mo agoif one of the advantages is making it copy-pastable then I would suggest the REXC viewer should give you the option to copy the REXC output, currently I have no way of knowing this by looking at your github or demo viewer another thing, I put in a 400KB json and the REXC is 250KB, cool, but ideally the viewer should also tell me the compressed sizes, because that same json is 65kb after zstd, no idea how well your REXC will compress edit: I think I figured out you can right click "copy as REXC" on the top object in the viewer to get an output, and compressed it, same document as my json compressed to 110kb, so this is not great... 2x the size of json after compression.
- creationix 7mo agoThanks for testing it out! Yes, the website could use some love to make everything more discoverable. The primary use case is not compression, it's just a nice side effect of the deduplication. This will never beat something like zstd, brotli, or even gzip. My production use cases are unique in that I can't afford the CPU to decompress to JSON and then parse to native objects. But with this format, I can use the text as-is with zero preprocessing and as a bonus my datasets are 18x smaller.
- creationix 7mo ago> 2x the size of json after compression Right and that makes sense. There is more information in here. The entire thing is length prefixed and even indexed for O(1) array lookups and O(log2 N) object lookups. If you don't care about random access and you don't mind the overhead of decompression, don't use RX.
- mpeg 7mo agoI think this makes sense, when you explain it like that, it might be a matter of cleaning up the docs a bit so the "why" of RX is more clear (admittedly, a README is not always the best channel for this!)
- creationix 7mo agoI've rewritten the framing in the README to first explain when you should use RX and when you should not. Most uses of JSON should probably stay JSON. Let me know what you think https://github.com/creationix/rx/blob/main/README.md#when-to-use-rx https://github.com/creationix/rx/blob/main/README.md#when-to...
- soco 7mo agoI have an idea, why don't we all go back using XML at this point, as any initial selling point / differentiator has been slowly eroded away?