8 ms·
If your dataset is large enough to benefit from such a hyperoptimized parser you might not benefit from the human readability anymore, which is the main reason
by stkdump 5y ago
If your dataset is large enough to benefit from such a hyperoptimized parser you might not benefit from the human readability anymore, which is the main reason for the required CPU cycles on the first place. So you should probably use a feature-equivalent binary format that is optimized for parsing speed. The only reason to use json then is that it is the lingua franca. So we as an industry should figure out which of the 100s of binary variants of json-similar formats has the same capabilities and unite around it.
- EwanToo 5y agoI largely agree, but most modern binary formats are more rigid with schemas and types. If you want the flexibility of JSON, I'm not sure you'll end up with something massively different from gzipped JSON
- lifthrasiir 5y agoJSON doesn't allow any custom type, so it is not "flexible" per se. Therefore you only need a format that supports the JSON data model and pretty much nothing else; CBOR [1] for example almost surely fits the bill. [1] https://cbor.io/ https://cbor.io/
- Waterluvian 5y agoThat’s an interesting thought. I am now very fascinated by what kind of data would be gigs in size but would need the flexibility of JSON.
- maccard 5y agoDatasets are not always provided by you; if your data source outputs json it doesn't matter why. Also just because your volume of data is measured in gigabytes, that doesn't mean it's a singular stream. As an example you could be handling tens of thousands of small requests.
- the8472 5y agoConverting a text format to a compact in-memory data structure takes extra CPU cycles. (de)compression takes extra cycles. The point of using a binary format is to achieve the same result while avoiding that overhead. For some data formats compression also has the downside that it prevents seeking.
- kevin_thibedeau 5y agoSerialized binary encodings don't require a schema. They can have all the expressability of JSON and more.
- dan-robertson 5y agoWell json is everywhere so I think it’s mostly too late for companies to change all their internal protocols and apis to not use it. And note that a big advantage of json is that many applications can process it without a schema—you don’t need to know that the username field is a certain length, type and at a particular place. It is also easier to know if json is valid than a random binary format. The advantage of something like the algorithm here is that you can put it into your shared libraries[1] to get a speed up for free. If you spend (making up numbers) 1% of your cpu time parsing json and this speeds things up 4x (note: possibly more because other parsers get more of an advantage in micro-benchmarks from the branch predictor) then if you have sufficiently many servers you can cut your server bill by 0.75% which can be a large amount of money for the largest companies. [1] this won’t work for all languages as the OP parses json into a flattish object where fields may be looked up rather than some language-specific data structures (like js objects or python dicts or whatever)
- m_mueller 5y agoyep. once you’ve had to deal with XMLs that come in without a schema alongside, you start appreciating getting a JSON. at least that properly distinguishes strings, nulls, integers and decimals.
- agent327 5y agoYour dataset could, instead of gigabytes in a single request, also be a very large number of small requests all sent to a server to handle. Each message might be easily human-readable, yet the whole system would benefit from expedient processing.
- secondcoming 5y agoYou could also potentially use plain CSV or TSV files. Not sexy, but they work.
- kozziollek 5y agoOh yes, CSV! Sure! Separated with commas or semicolons? With quotes or without? What decimal separator? No thanks. Use Protobufs, Parquet, Avro, etc.
- cogman10 5y agoI've gotta say, I don't really understand why the industry is so in love with human readability for computer to computer interaction. It seems like a notion that started in the early internet and just refuses to die. Everything on the internet is minified and compressed at this point, so the entire idea of "human readability" left the station years ago. Yet I'll still see a primary complaint against the likes of http2+ being "it's a binary format! On NO!". JSON seems like a similar relic. We use it not because it's fast, but because we like the idea that it's easy to decode (Even if in practice that almost never happens).
- AYBABTME 5y agoIt's the same idea as designing hardware for repairability. Sure you can repair something that's welded on, but it's much easier if it's bolted through instead.
- cogman10 5y agoIt's not the same. You are almost certainly passing these through tools (even built into the browser) that are doing extra processing to make it more readable. Let's assume, for example, CBOR ends up taking over JSON. Do you not think browsers wouldn't have a CBOR parser to make it more readable? This isn't bolt vs welding, this is bolt vs bolt with a washer. Yes there's a small extra step, but not some sort of insurmountable hurdle.
- pixl97 5y agoThe questions about readability when things go wrong. When binary data breaks, mostly your just screwed in figuring out went wrong, especially when it causes your decoder to fall over. When ASCII data breaks often a view of the data can point at something even when tools can't.
- rytcio 5y agoReading a binary file is no different than reading ascii. I experimented with writing a debugger, which meant learning how to parse executable ELF binaries. It really wasn't bad at all. It just takes a bit more time to figure out what the bytes/data are expect at what offsets. The real un-said reason is a) programmers are lazy b) its easier/cheaper to hire people who can read ascii than find someone who is a bit more experienced and can figure out how to debug binary parsing.
- Zababa 5y ago> If your dataset is large enough to benefit from such a hyperoptimized parser I don't understand that point. Is there an overhead in starting the parsing, which makes regular parsing faster unless you have a large JSON file? If not, why wouldn't you want faster JSON parsing?
- lumost 5y agoMost languages have a built in parser these days, pulling in a c dep can be painful in many build tools and languages. The csimdjson api is sufficiently different to not as idiomatic. Most language json parsers are very fast already. Personally I’m using csimdjson in a project with 100s of TB of json to burn through. This data should not be json formatted, but migrating away from json would require modifications to hundreds of different systems.
- Zababa 5y agoThat's a fair point. Though this could probably be used in interpreters/JIT/C++ projects, which is already a lot. And this gives a good template on how to optimize JSON parsing for other projects.
- SigmundA 5y agoText is a binary format that just happens to have decoders (ASCII, UTF8 etc) everywhere in order to parse and display it to humans along with more specific parsers for things like JSON. I wish we could all settle on a binary structured format thats not limited to text encoding and more like JSON (key value) that basically everyone uses with real data types (efficiently encoded numbers, dates, binary , etc). Text files would just be something like { contentType: "text/plain" content : "text here" } while an image could be { contentType: "image/jpeg", content: <real binary data not base64> } you could also add whatever other metadata you want and all of it is retained and easily parseable. Other more structured formats obviously wouldn't just be a blob for content and every system out there would have a structured binary format viewer/editor just like there are text viewers and editors now. Sqlite is kinda used this way and has some nice properties like indexes and transactions but is relational instead of hierarchal which can be good and bad, not as straight forward to just view or navigate. Protobuf is another used quite a bit now, honestly I don't care just something everyone agree upon that can encode more structure efficiently but still can be easily inspected everywhere. Doubt this will any time soon but it does seem inevitable in the long run that we figure out a way to send data between system and what a string, number, date etc is and stop having text encoding and escaping issues.
- kortex 5y agoI don't know why the ASCII characters for record and field separators don't get more love. That is what they are there for. My suspicion is because they aren't type-able, so people rarely encounter them.
- SigmundA 5y agoStill have the issue of binary data encoding as well since the data will have those bytes in it so they need to be escaped vs a format that defines how to encoding binary as is efficiently (length delimited).
- legulere 5y agoYou mean like CBOR? https://en.wikipedia.org/wiki/CBOR https://en.wikipedia.org/wiki/CBOR
- orasis 5y agoprotocols have two sides - a producer and a consumer. The consumer often doesn’t have control of the producer, especially in the case of analytics.
- pickledish 5y agoI guess an example would be along the lines of protobufs, e.g. Cap’n Proto https://capnproto.org/ https://capnproto.org/ (which I have never used but really enjoy just for its charming website design alone)
- Epa095 5y agoIf the improvements scale, you can still get benefit for smaller use cases. And even if each json is small, if you have millions of them the tiny improvements add up. We parse quite a bit of json, we fetch and store json regularly from external api's. We can't ask for a different format, and for each of the individual requests json makes sense, the response is only a few hundred kbs. But over time there is a lot of data. We convert them to parquet, but each json needs to be read at least once.
- tfsh 5y agoProtobufs would be a good contender here
- asdfge4drg 5y agouh yes that would work perfectly. https://xkcd.com/927/ https://xkcd.com/927/
- ec109685 5y agoWhy do you think human readable adds appreciable overhead? If you want to create a flexible interchange format, that is going to require some sort of parse step whether the format is text or binary. That said, probably FlatBuffer would be even an optimized json parser, but json is crazy fast if the parser takes advantage of all modern processor optimizations.
- koolba 5y agoThere’s no way reading a series of text is going to beat a machine native byte representation, particularly for things like range limited integers. If you plan things right with proper word alignment, you don’t even have to copy the data to process it as you can directly dereference the byte offsets as machine native ints.
- ec109685 5y agoBut that isn’t an interchange format given different processors represent such things differently.
- wilde 5y agoHow do you debug? Human readability is valuable regardless of the size of the dataset.
- rixed 5y ago"we as an industry" hurts. It's funny how I feel questioning JSON here on HN is like starting a discussion about politics in a family diner. I'm forever grateful to Google for demonstrating that all "the industry" is not fully committed to HTTP, JSON and scripting languages. It's honestly a relief to have encountered some sanity somewhere. Oh god please help me, I just did it, I questioned JSON on HN!