7 ms·
Deserializing JSON Fast (2020)
- dundarious 5y agoGoing further down the road of only doing what your requirements demand (not writing a general purpose parser), I really enjoyed this talk. https://media.handmade-seattle.com/context-is-everything/ https://media.handmade-seattle.com/context-is-everything/
- lifthrasiir 5y agoDon't get me wrong, I mostly agree to what the talk wants to say, but the talk itself doesn't do a good job at giving convincing reasons. In fact I think the OP is a much better replacement for that talk. There are several factual inaccuracies in the talk. For example on-demand or in-situ parsing has been a thing especially for many high-performance C++ JSON parsers, and they are exactly designed for the constraints described in the talk (no needs for parse tree, in-place mutation or cross referencing). So everything before the specialized integer array parsing can be mostly automated without heavy hammers like SIMD---those parsers would likely use that, but I don't care. You still need to recognizes those constraints in the first place but you don't always need to write your own code for making use of them. He did mention that dependencies are or can be liabilities, yet failed to state that more codes and contexts are also liabilities. It is correct to recognize implicit contexts that may be useful for the optimization, but it doesn't mean that every context should be utilized. For example later examples assume x86(-64), which is increasingly becoming a bad decision in general thanks to ARM servers. It is a balancing act to choose which context to utilize and which to ignore (so that the code can remain more generic for later uses or easier to understand) and the talk is mostly silent about this aspect. He also fails to mention testing, which is the main reason I believe the OP is a better substitute for the talk. You can't exactly compare your shiny optimized code with an older version of it, because you have distilled implicit contexts into the new code so they can be functionally different. You need to somehow make those implicit contexts into explicit contracts, probably using some other codes that are not continuously run, documentations (which everyone seem to overlook!) or folk, eh, collective team knowledge. This is a hard problem by its own and again the talk is pretty much oblivious about that.
- dundarious 5y agoI think the talk is very carefully laid out as an incremental journey, and each stepping stone involves contextual decision-making. I don't think Andreas is saying "you must end up with the SSE2 implementation at the end", or even that approaches like the OP are illegitimate -- I'm commending the OP for thinking about the problem the right way! But using machine-specific intrinsics is just another dependency decision very similar to deciding to use a given library -- you may choose otherwise, for good reason. I would have loved the talk and probably still thought of it and posted it, even if it ended before the intrinsics (but I think he does an excellent job at that part too). But porting SSE2 to Neon is actually pretty easy -- if you use https://github.com/DLTcollab/sse2neon https://github.com/DLTcollab/sse2neon, IME it's very easy to do incrementally (or avoid or postpone indefinitely, depending on your needs).
- olliej 5y agoSo this isn't really deserializing JSON as some of the optimizations they are making are based on them actually targeting a JSON subset. That removes a bunch of conditions and branches. Still, it's interesting, and a great example of ways handling trusted data is faster than untrusted (and I guess more importantly how people writing handlers for untrusted code can accidentally "optimize" things into security flaws :D) That said, in my experience the real cost of JSON parsing turns out to be constructing the in memory representation of the sources. I'm sure JSON parsing for the purpose of deserializing specific types would be faster as you might be able to avoid some allocations (e.g. imagine a vector type `struct Vector { float elements[4]; }` a generic JSON representation would result in two allocations)
- beached_whale 5y agoIn my experience, doing this in C++, disabling checks is maybe 5-20%, more often around 10%, not really worth doing outside benchmarks. When it's already going at GB/s, just do the checks and use the types to get more checks. Allocations, have been a huge win. Parsing to a vector in many of the JSON benchmarks benefitted from reserving the memory up front, even if not known. In testing the cost of them, it was around 50% of the bench time. Using a bump allocator to test doubled the perf when I did it.
- olliej 5y agoOh they have other things: no white space is a big one. Depending on what they’re allowing they may be able to stop handling character escapes, etc But actual allocation ends up being killer in the JSC parser. I even tried using structure caching to speed up object creation, but it simply did not seem to matter: either there’s too little data to amortize allocations for the necessary data structures, or you’re parsing enough that simple allocations and GC start to happen (and code in c++ can’t optimize the allocations as much as the JIT can)
- beached_whale 5y agoI forget what I got when I benchmarked assuming minified, but I already optimized on that anyways. SIMDJson burns through pretty JSON but less so with minified. With C++, there are generally just less allocations, but they probably cost more, not always and I haven't measured malloc/new for small things where they often have a pool setup before going to mmap. They still lock though. But one can often just guess when the number isn't known without much cost on most systems. Choose a size like 1kb and remove the smaller up front allocations, or if the size is known reserve that.
- beached_whale 5y agoNot Rust, but for C++ I have found a lot of memory and CPU perf can be realized by knowing what is being parsed into and using that. Most JSON parsing optimize on array/object/number/bool/null and not things like unsigned integer being much faster to parse than a float(floating point hit cpu bottlenecks around 1GB/s). Additionally, allowing for compile time options to further optimize is much easier. Essentially most JSON parsers are type erased. https://github.com/beached/daw_json_link https://github.com/beached/daw_json_link So this correlates with some of the findings the post has. One optimization that should show great promise is that if one knows that they have an array of some types like numbers/unsigned int/strings, that can greatly simplify and allow for SIMD'ification of it.
- leni536 5y agoCheck out https://github.com/beached/daw_json_link https://github.com/beached/daw_json_link , it provides a non-typeerased way to parse JSON straight into user-defined data structures.
- beached_whale 5y agoIm the author :)
- leni536 5y agoOh, lol, didn't check the name there, and didn't through the whole comment either.
- staticassertion 5y ago> The upside of those tradeoffs are that it improves our query throughput by about 20% compared to simd_json Woah, that's impressive. I wonder if `str` is maybe an antipattern for a lot of JSON use cases where you can perform utf8 validation lazily or avoid it altogether for data you don't need. Given a service architecture with mTLS, and so long as you aren't doing anything sensitive based on the data, I could see an approach like this being valuable. That said, I also wonder if it's worth pushing JSON into use cases like this to begin with?
- chrismorgan 5y agoBoth squirrel-json and simd_json are doing UTF-8 validation, that’s not the difference. squirrel-json does it as a first pass before the parsing, simd_json does it during parsing, but both are doing it.
- Youden 5y agoNot that the optimisation isn't interesting but if the query is column-based, wouldn't it make more sense to use a binary column-oriented format like Apache Parquet [0]? Or perhaps a binary format like Protobuf, Webpack, Cap'n'proto etc.? Are there other concerns that cause JSON to make sense? [0]: https://parquet.apache.org/documentation/latest/ https://parquet.apache.org/documentation/latest/
- BerislavLopac 5y agoOur Amazon Ion [0], which is compatible with JSON. [0] https://amzn.github.io/ion-docs/ https://amzn.github.io/ion-docs/
- tomnipotent 5y agoSeq seems to be targeting application logs, where JSON (and jsonl) is a common format and it makes sense.
- jsnell 5y agoThere's three data formats involved: 1. The one used for ingesting data. 2. The one used for storage. 3. The one used for exporting data. There is no reason for all three formats to be the same, and there are good reason for them not to be: the actual use cases are requiring very different properties from the data format. Since you'd expect the average log record to be ingested once, processed a gazillion times, and returned as output <<1 times, any data conversion costs would get amortized over a lot of processing. Even minor efficiency gains in the at-rest storage format should pay off quickly. Your answer is addressing format 1: the clients are using JSON anyway, so it makes sense for the ingestion format to be JSON. That's making the bet that any data conversion costs can't be paid back over the lifetime of the data. But they are already not using the raw client input as the storage format! At a minimum they're validating / minifying / canonicalizing the JSON, so they're already paying the cost of doing a format conversion. (But all of this is obvious, and clearly the people writing this service are smart and would have thought of it. So it seems like there's a non-obvious answer for using minified JSON for storage at rest.)
- jheriko 5y ago"we don't need error handling because we output the thing" i'm wondering why the hell its not a faster format? json is awful for storing ... anything, but a good compromise given that "everyone knows what it is"
- beebmam 5y agoIf you want to optimize for minimal deserialization time, consider using FlatBuffers for serialization: https://google.github.io/flatbuffers/ https://google.github.io/flatbuffers/
- ra-mos 5y agoFlatbuffers have been good in a demux/parallel architecture I’ve been managing. We split single requests into ~100-500 requests — each using the same Flatbuffer object (list of items). Each sub request reads header data to pull some percentage (1/100-500) of data out of the flatbuffer it needs to process. The library is still fairly low level, and our abstraction has become complex. Still, the best performance approach we’ve been able to find.
- k__ 5y agoHow do these binary protocols work in the browser? I read once, a custom library written in JS will probably never be as fast as the JSON parser shipped in browsers.
- tyingq 5y agoI don't see how it's possible that "look at every single byte for the unescaped balanced quote/bracket that ends the field/record" could be faster than "read the length, read the whole field in optimal buffer sizes". I suppose there's other stuff that's being done that could be suboptimal, but that's the key difference.
- inglor 5y agoParsing JSON is tricky, you need to figure out the structure as you go. JavaScript implementations can reuse their JIT machinery to figure out the structure of parsed JSON (here's shapes if you're not familiar https://mathiasbynens.be/notes/shapes-ics https://mathiasbynens.be/notes/shapes-ics ) You _can_ do it outside the engine but it'd be tricky and you'd have to either write a lot of code _or_ to trick the engine to reuse the JIT machinery for your parser. Protocols with schemas are much much easier to parse efficiently so I suspect it'd be as fast in library written JS.
- 5y ago