6 ms·
RX – a new random-access JSON alternative
- NoSalt 7mo agoWhy do we need an "alternative" when JSON, itself, is so fantastic?
- creationix 7mo agothe project framing needs some help perhaps. JSON is really good at a lot of use cases that this will never replace. But there are cases where JSON is currently used where this is much better. In particular large unstructured datasets where you only need to read a tiny subset of the data in a single request. Maybe a better framing would be no-sql sqlite?
- creationix 7mo agoA new random-access JSON alternative from the creator of nvm.sh, luvit.io, and js-git.
- barishnamazov 7mo agoYou shouldn't be using JSON for things that'd have performance implications.
- hrmtst93837 7mo ago[flagged]
- creationix 7mo agoYep. I did try binary formats first. I tried existing ones like CBOR, I tried making my own like Nibs. The text encoding is an operational concern, not a technical one. This is the same reason I've been advocating for JSONL at work. It's not ideal technically, but it's a good balance of technically good enough while being also human friendly when things go wrong. - https://vercel.com/blog/how-we-made-global-routing-faster-with-bloom-filters https://vercel.com/blog/how-we-made-global-routing-faster-wi... - https://vercel.com/blog/scaling-redirects-to-infinity-on-vercel https://vercel.com/blog/scaling-redirects-to-infinity-on-ver... RX is one step towards less human friendly, but more machine friendly. I try to keep things balanced in my designs.
- Spivak 7mo agoCan you imagine if a service as chatty and performance sensitive as Discord used JSON for their entire API surface?
- creationix 7mo agoAs with most things in engineering, it depends. There are real logistical costs to using binary formats. This format is almost compact as a binary format while still retaining all the nice qualities of being an ASCII friendly encoding (you can embed it anywhere strings are allowed, including copy-paste workflows) Think of it as a hybrid between JSON, SQLite, and generic compression. This format really excels for use cases where large read-only build artifacts are queried by worker nodes like an embedded database.
- Asmod4n 7mo agoThe cost of using a textual format is that floats become so slow to parse, that it’s a factor of over 14 times slower than parsing a normal integer. Even with the fastest simd algos we have right now.
- meehai 7mo agoand with little data (i.e. <10Mb), this matters much less than accessibility and easy understanding of the data using a simple text editor or jq in the terminal + some filters.
- xxs 7mo agowhat do you mean by little data, most communication protocols are not one off
- creationix 7mo agoAlso good luck parsing 10 MiB of JSON in a loop that can't tolerate blocking the CPU for more than 10ms. What's expensive is very relative to the use case.
- HelloNurse 7mo agoSo it depends. Float parsing performance is only a problem if you parse many floats, and lazy access might reduce work significantly (or add overhead: it depends).
- squirrellous 7mo agoI agree in principle. However JSON tooling has also got so good that other formats, when not optimized and held correctly, can be worse than JSON. For example IME stock protocol buffers can be worse than a well optimized JSON library (as much as it pains me to say this).
- tabwidth 7mo agoYeah the raw parse speed comparison is almost a red herring at this point. The real cost with JSON is when you have a 200MB manifest or build artifact and you need exactly two fields out of it. You're still loading the whole thing into memory, building the full object graph, and GC gets to clean all of it up after. That's the part where something like RX with selective access actually matters. Parse speed benchmarks don't capture that at all.
- xxs 7mo agoas parser: keep only indexes to the original file (input), dont copy strings or parse numbers at all (unless the strings fit in the index width, e.g. 32bit) That would make parsing faster and there will be very little in terms on tree (json can't really contain full blow graphs) but it's rather complicated, and it will require hashing to allow navigation, though.
- creationix 7mo agoyep. I built custom JSON parsers as a first solution. The problem is you can't get away from scanning at least half the document bytes on average. With RX and other truly random-access formats you could even optimize to the point of not even fetching the whole document. You could grab chunks from a remote server using HTTP range requests and cache locally in fixed-width blocks. With JSON you must start at the front and read byte-by-byte till you find all the data you're looking for. Smart parsers can help a lot to reduce heap allocations, but you can't skip the state machine scan.
- magicalhippo 7mo ago> The real cost with JSON is when you have a 200MB manifest or build artifact and you need exactly two fields out of it. There are SAX-like JSON libraries out there, and several of them work with a preallocated buffer or similar streaming interface, so you could stream the file and pick out the two fields as they come along.
- Levitating 7mo agoJSON is human-readable, why even compare it with this. Is any serialization format now just a "JSON alternative"?
- dietr1ch 7mo agocat file.whatever | whatever2json | jq ? (Or to avoid using cat to read, whatever2json file.whatever | jq)
- creationix 7mo agoOr in this case, just do `rx file.rx` It has jq like queries built in and supports inputs with either rx or json. Also if you prefer jq, you can do `rx file.rx | jq`
- dietr1ch 7mo agowow, on that case then using `jq` is just a presentation preference at the very last step unless jq is more expressive (which might be the case given how long it has been around?).
- creationix 7mo agoright, the jq query language is much more complex and featureful than the simple selector syntax I added to the rx-cli. But more could be added later as needed or it could just stream JSON output. It would be pretty trivial to hook up a streaming JSON encoder to rx-cli which could then pipe to jq for low-latency lookups. The problem is jq would need to JSON parse all that data which will be expensive.
- Gormo 7mo agoThat's not really random access, though. You're effectively just searching through the entire dataset for every targeted read you're after. What might be interesting is to have a tool that processes full JSON data and creates a b-tree index on specified keys. Then you could run searches against the index that return byte offsets you can use for actual random access on the original JSON. OTOH, this is basically just recreating a database, just using raw JSON as its storage format.
- Spivak 7mo agoI love these projects, I hope one of them someday emerges as the winner because (as it motivates all these libraries' authors) there's so much low hanging fruit and free wins changing the line format for JSON but keeping the "Good Parts" like the dead simple generic typing. XML has EXI (Efficient XML Interchange) for precisely the reason of getting wins over the wire but keeping the nice human readable format at the ends.
- snthpy 7mo agoTIL. EXI looks useful. Now I just wish there was a renderer in the pugjs format as I find that terse format much pure readable than verbose XML. I also find indentation based syntax easier to visually parse hierarchical structure.
- garrettjoecox 7mo agoVery cool stuff! This did catch my eye, however: https://github.com/creationix/rx?tab=readme-ov-file#proxy-behavior https://github.com/creationix/rx?tab=readme-ov-file#proxy-be... While this is a neat feature, this means it is not in fact a drop in replacement for JSON.parse, as you will be breaking any code that relies on the that result being a mutable object.
- creationix 7mo agoTrue, the particular use case where this really shines is large datasets where typical usage is to read a tiny part of it. Also there is no reason you couldn't write an rx parser that creates normal mutable objects. It could even be a hybrid one that is lazy parsed till you want to turn it mutable and then does a normal parse to normal objects after that point.
- btown 7mo agoThis is really interesting. At first glance, I was tempted to say "why not just use sqlite with JSON fields as the transfer format?" But everything about that would be heavier-weight in every possible way - and if I'm reading things right, this handles nested data that might itself be massive. This is really elegant. My one eyebrow raise is - is there no binary format specification? https://github.com/creationix/rx/blob/main/rx.ts#L1109 https://github.com/creationix/rx/blob/main/rx.ts#L1109 is pretty well commented, but you can't call it a JSON alternative without having some kind of equivalent to https://www.json.org/ https://www.json.org/ in all its flowchart glory!
- creationix 7mo agoThanks. I had this for older versions, but forgot to write it up again for the latest version. One old version that is meant to be more human readable/writable is jsonito https://github.com/creationix/jsonito https://github.com/creationix/jsonito I'll add similar diagrams and docs for the format itself here.
- creationix 7mo agoInitial format docs are now here: https://github.com/creationix/rx/blob/main/docs/rx-format.md https://github.com/creationix/rx/blob/main/docs/rx-format.md Railroad diagrams will come later when I have more time.
- btown 7mo agoNeat! In case you took me too literally: railroad diagrams are fun, but far from the only way to give spec level clarity, so don’t feel you need to overindex on my silly comment! I am curious why it’s parsed right to left. Is this so that you could add new data to a top-level JSONL-esque list, solely by rewriting the end of the data structure, and not needing to change the beginning (or worst-case shift every single byte of data, if you need a longer count)? It’s an interesting design tradeoff, because you can’t show a partial parse if you’re streaming the content naively beginning to end, which is a bit odd in a world where streams that begin to render token-by-token are all the rage. But if you have an ability to do range queries, it’s quite effective, and it does allow for those incremental updates!
- benatkin 7mo agoInteresting. I've heard about cursors in reference to a Rust library that was mentioned as being similar to protobuf and cap'n proto. Does this duplicate the name of keys? Say if you have a thousand plain objects in an array, each with a "version" key, would the string "version" be duplicated a thousand times? Another project a lot of people aren't aware of even though they've benefitted from it indirectly is the binary format for OpenStreetMap. It allows reading the data without loading a lot of it into memory, and is a lot faster than using sqlite would be. Edit: the rust library I remember may have been https://rkyv.org/ https://rkyv.org/
- creationix 7mo ago> Does this duplicate the name of keys? Yes, the format allows for objects to be stored with a pointer to a shared schema (either an array of keys or another object that has the desired keys) The current implementation is pretty close to ideal when deciding to use this encoding.
- Shahbazay0719 7mo ago[flagged]
- dtech 7mo agoIt's not quite clear to me why you'd use this over something more established such as protobuf, thrift, flatbuffers, cap n proto etc.
- maxmcd 7mo agoThose care about quickly sending compact messages over the network, but most of them do not create a sparse in-memory representation that you can read on the fly. Especially in javascript. This lib keeps the compact representation at runtime and lets you read it without putting all the entities on the heap. Cool!
- IshKebab 7mo agoAmazon Ion has some support for this - items are length-prefixed so you can skip over them easily. It falls down if you have e.g. an array of 1 million small items, because you still need to skip over 999999 items to get to the last one. It looks like RX adds some support for indexes to improve that. I was in this situation where we needed to sparsely read huge JSON files. In the end we just switched to SQLite which handles all that perfectly. I'd probably still use it over RX, even though there's a somewhat awkward impedance mismatch between SQL and structs.
- creationix 7mo agoI did seriously consider SQLite, but my existing datasets don't map easily to relational database tables. This is essentially no-sql for sqlite.
- creationix 7mo agoExactly. Low heap allocations when reading values is one of the main driving factors in this design!
- konart 7mo agoWhat if you are reading from a service which already have an established API? It's not like you can just tell them to move to protobuf.
- WatchDog 7mo agoCool project. The viewer is cool, took me a while to find the link to it though, maybe add a link in the readme next to the screenshot.
- transfire 7mo agoI am a little confused. Is this still JSON? Is it “binary“ JSON?
- SV_BubbleTime 7mo agoIt’s neither! Sample output: 'fdiscovered,aextreme,7danger,6+1A+16;6level_range,b:QThe Heap ,d'th Human unreadable, ascii output. Line up and get yours today!
- creationix 7mo agoit's not really possible to stay human readable and get the compression levels and random access properties I was going for. But it is as human tooling friendly as possible given the constraints.
- SV_BubbleTime 7mo ago>it's not really possible I find it obvious that your first attempt failed. Try again, you have not even remotely failed enough if you are making the argument that this is kinda readable. Yes, ascii words are easy to pick out, you didn’t do that, you did the part that makes it all harder.
- _flux 7mo agoIt doesn't seem the actual serialization format is specified? Other than in the code that is. Is it versioned? Or does it need to be..
- 50lo 7mo agoThe biggest challenge for formats like this is usually tooling. JSON won largely because: every language supports it, every tool understands it. Even a technically superior format struggles without that ecosystem.
- latexr 7mo agoAnd that in turn affects tool adoption. I have dabbled in Lua for interacting with other software such as mpv, but never got much into the weeds with it because it lacks native JSON support, and I need to interact with JSON all the time.
- creationix 7mo agoyeah, LuaJIT is one of the use cases I had in mind working on this. JSON is pretty fast in modern JS engines, but in Lua land, JSON kinda sucks and doesn't really match the language without using virtual tables. JSON has `null` values with string keyds, but lua doesn't have `null`. It has `nil`, but you can't have a key with a nil value. Setting nil deletes the key Lua tables are unordered. But JS and JSON are often ordered and order often matters. RX, however matches Lua/LuaJIT extremely well and should out-perform the JS Proxy based decoder using metatables. Since it's using metatables anyway do to the lazy parsing, it's trivial to do things like preserve order when calling `pairs` and `ipairs` and even including keys with associated null values. You can round trip safely in Lua, which is not easy with most JSON implementations.
- StephenZ15ga67 7mo ago[flagged]
- jbverschoor 7mo agoSo this is two things? A BSON-like encoding + something similar to implementing random access / tree walker using streaming JSON? Docs are super unclear.
- pshirshov 7mo agoLooks similar to https://github.com/7mind/sick https://github.com/7mind/sick
- creationix 7mo agoYou're right. Some important differences: sick is binary, rx is textual (this matters for tooling) sick has size limits (65534 max keys for example. I have real-world rx datasets reaching this size already) rx uses arbitrary precision variable-length b64 integers. There are no size limits anywhere inherit in the format, just in implementations. sick does not preserve object key order rx preserves object key order, but still implements O(log2 N) lookups for object keys. etc.
- openclaw01 7mo ago[dead]
- derodero24 7mo ago[flagged]
- gritzko 7mo agoI recently created my own low-overhead binary JSON cause I did not like Mongo's BSON (too hacky, not mergeable). It took me half a day maybe, including the spec, thanks Claude. First, implemented the critical feature I actually need, then made all the other decisions in the least-surprising way. At this point, probably, we have to think how to classify all the "JSON alternatives" cause it gets difficult to remember them all. Is RX a subset, a superset or bijective to JSON? https://github.com/gritzko/librdx/tree/master/json https://github.com/gritzko/librdx/tree/master/json
- deleted 7mo ago[deleted]
- SV_BubbleTime 7mo agoYou went from BSON to your own and skipped CBOR and Protobuf? … I wonder if you would have made different decisions without Claude vibing you in a direction?
- creationix 7mo agoThe current format version is the exact same feature set as JSON. I even encode numbers as arbitrary precision decimals (which JSON also does). This is quite different from CBOR which stores floats in binary as powers of 2. I could technically add binary to the format, but then it would lose the nice copy-paste property. But with the byte-aware length prefixes, it would just work otherwise.
- bsimpson 7mo agoIt feels petty to show up with a naming not, but the name is unfortunately/confusingly similar to the already well-known RxJS. Why is it called RX?
- creationix 7mo agoI'm happy to hear suggestions. This format was actually the internal .rexc bytecode for Rex (routing expressions), but when I realized it was actually a pretty good standalone format, I renamed it `.rx` for short. I am aware of RxJS, but I think that `rx-format` is different enough and `.rx` file extensions are unique enough, it's not too confusing.
- AliEveryHour16 7mo ago[dead]
- TKAB 7mo agocould this be useful for embedding info in server generated web pages that are then picked up by a JavaScript. e.g. a tom-select country picker that gets its data from an embedded RX structure?
- creationix 7mo agoyes, this would work very well for any case where you have embedded databases of unstructured data that you want to query in a website or edge server
- dietr1ch 7mo agoA tiny note on the speed comparison: The 23,000x faster single-key lookup seems a bit misleading to me. Once you get the computational complexity advantage, then you can make it as much times faster as you want. In these cases small instances matter to judge constants, and to the average (mean?) user, mean instance sizes. I'm not sure how to sell the advantage succinctly though. Maybe just focus on "real-world" scenarios, but there's no footnote with details on the comparison
- creationix 7mo agoThat benchmark is a fair comparison for a real-world production workload and use case. Sadly I can't share the details. But suffice it to say that the dataset is a huge object with tens of thousands of paths as keys and moderately large objects as values (averaging around 3KB of JSON each) all with slightly different shapes. The use is reading just a few entries by path an then looking up some properties within those entries. The benchmark (or is supposed to) measures end-to-end parse + lookup. JSON: 92 MB RX: 5.1 MB Request-path lookup: ~47,000x faster Time to decode a manifest and look up one URL path: JSON: 69 ms REXC: 0.003 ms Heap allocations: 2.6 million vs. 1 JSON: 2,598,384 REXC: 1 (the returned string)
- killbot5000 7mo agoThe documentation reference a “decode” function, and it’s imported to the example code, but it’s never called. I’m not sure what the API is after reading the examples.
- DaleBiagio 7mo agoJSON's dominance is one of the most accidental success stories in computing. Douglas Crockford didn't design it — he said he "discovered" it. It was already there in JavaScript's object literal syntax, which itself traces back to Brendan Eich's 10-day sprint in 1995. A data format that conquered the internet was a side effect of a language built under absurd time pressure. Every attempt to replace it has to overcome that kind of accidental ubiquity, which is much harder than overcoming a technical limitation.