26 ms·
Parsing JSON Is a Minefield (2018)
- Arrezz 7y agoThe bigger question is, what is there to be done? What is the road to a more uniform handling of JSON? I've handled some JSON before and it's usually fairly easy untill you catch one of these strange implementation quirks. But I'm not sure that those quirks can be ironed out at this point.
- hirundo 7y ago> The bigger question is, what is there to be done? Something like https://github.com/nst/JSONTestSuite https://github.com/nst/JSONTestSuite? Could a parser test suite be an official component of an RFC like 8259?
- erikpukinskis 7y agoCan you help me understand the problem? These things seem like corner cases that you could just Not Do(TM) and then you don’t have to worry about it. What am I missing, when do these gotchas become an actual problem for you as a developer? I’m not sure I’ve used any technology that was free of footguns, and JSON appears to have fewer of them than the average programming language or library.
- falcolas 7y agoI call this kind of answer “The C Answer”. “Who cares if this particular combination of code results in undefined behavior? Just don’t do it!” > when do these gotchas become an actual problem for you as a developer? Whenever you have to deal with JSON produced by “not you”, or when you have to deal with JSON that may have been corrupted in some fashion along the way.
- erikpukinskis 7y agoI’m probably just not used to pure “all behavior is defined” systems, so I appreciate your perspective. What industry do you work in? I don’t see many systems like that in my industry. I work like hell to push things in that direction, but it’s a best case of “we went from 5% well defined behavior to 50%” after many years of effort.
- deleted 7y ago[deleted]
- umvi 7y agoI'm fine with a parser that doesn't get all of the corner cases as long as it fails gracefully. Really, the only time it would matter is if you are parsing user-provided JSON and said user was trying to exploit your parser somehow. But 99% of the time, I'm not parsing user-provided JSON, so I don't ever encounter these corner cases and parsing/serialization works great.
- majewsky 7y ago> 99% of the time, I'm not parsing user-provided JSON I take it you've never implemented a service with a REST API.
- umvi 7y agoI have, but I work
- umvi 7y agoOops, somehow the second half of the comment got truncated. I meant to say: "I have, but I work mainly with embedded systems, and so there aren't usually POST or PUT APIs, just GET. So JSON is nearly always generated by the device and consumed by whoever is querying the device, not the other way around."
- deleted 7y ago[deleted]
- juliusmusseau 7y agoWhat about the 2^63 corner-case? Consider this JSON: {"key": 9223372036854775807}. With most parsers it never fails. But... some JSON parsers (include JS.eval) parse it to 9223372036854776000 and continue on their merry way. The problem isn't user-provided JSON here. The problem is user-provided data (or computer-provided data) that's inside the JSON. rachelbythebay's take (http://rachelbythebay.com/w/2019/07/21/reliability/ http://rachelbythebay.com/w/2019/07/21/reliability/): On the other hand, if you only need 53 bits of your 64 bit numbers, and enjoy blowing CPU on ridiculously inefficient marshaling and unmarshaling steps, hey, it's your funeral.
- exabrial 7y agoI miss the days of strongly typed schemas. It's much easier to fail gracefully.
- Spivak 7y agoThrow your support behind https://json-schema.org https://json-schema.org it's a great effort.
- rapsey 7y agoOr use something sane like protocol buffers
- bcrosby95 7y agoHow would I pad my resume if I didn't keep reinventing the same wheels over and over?
- Pfhreak 7y agoI can understand if protocol buffers had some technical issue that you disagreed with. Or if you had a preference for a different solution because of some reason. But this comment seems needlessly cynical and doesn't actually offer any rebuttal to the parent's point. I find technical discussions of these sorts of things interesting and a great way for new people to learn about the tradeoffs, maybe you could offer a reasoned opinion on why not use protocol buffers?
- yongjik 7y agoI think GP is making fun of people who do not use protobufs and instead try newer, less proven ones...
- jakear 7y agoA problem we're finding with those is the massive runtime dependency you get trying to include protobuf in browser.
- truth_seeker 7y agoJust recently V8 the JS engine rewrote their JSON parsing code to achieve upto 2.7x faster parsing and also making it memory efficient. Ref link:- https://v8.dev/blog/v8-release-76 https://v8.dev/blog/v8-release-76
- majewsky 7y agoAh, so that's the source of that Chrome bug that we saw last week. Customers on Chrome for Windows (only that, not Chrome for Linux or macOS) were complaining that the search on our statically-generated documentation site was not working. The search is implemented by a JavaScript file that downloads a JSON containing a search index, and it turns out that this search index had too much nesting for Chrome on Windows's JSON parser. This would reliably produce a stack overflow: JSON.parse(Array(3000).join('[')+Array(3000).join(']')) We were about to report a bug when we noticed that the problem was fixed in Chrome 76, and the users in question were still on Chrome 75.
- zazagura 7y agoPretty weird for a JSON parser to be platform dependent.
- CamouflagedKiwi 7y agoMaybe the stack is shallower on Windows, or the calling convention takes more space per function and is enough to push it over the edge.
- nh2 7y agoRecursion based implementation of parsers in languages with limited stack size is a programmer mistake.
- justincormack 7y agoAll languages have limited stack size.
- saagarjha 7y ago> For example, Xcode itself will crash when opening a .json file made the character [ repeated 10000 times, most probably because the JSON syntax highlighter does not implement a depth limit. FWIW, this appears to have been fixed recently.
- tonyedgecombe 7y agoSlightly off topic but Xcode crashing has been very common for me, as much as I like macos the tooling leaves me missing Windows development.
- trilila 7y agoPython, go, rust and other obscure or archaic languages struggle with json. I recommend using a modern language such as js/nodejs, as it is meant for the web - where json is king among formats.
- vageli 7y ago> Python, go, rust and other obscure or archaic languages struggle with json. I recommend using a modern language such as js/nodejs, as it is meant for the web - where json is king among formats. In what ways are the languages you listed obscure or archaic?
- TylerE 7y agoThey don't have leftpad.
- vorticalbox 7y agoin js/node is is a valid object { hello: 'world' } in python its not valid json = { "hello": "world" } in python you also can't do json.hello like in js, it would be json['hello']
- jkern 7y agoI'm going to be generous and hope this is some really dry humor and not an actual comment
- aflag 7y agoI don't understand what you're saying {hello: 'world'} is not valid json and neither Python's nor javascript's json parser from the standard library parse that.
- coldtea 7y agoNone of those mean Python is "obscure or archaic" or "unfit for json" etc. Regarding what is a valid object, it's about syntax. Python has a different syntax than JS. That doesn't mean Python is worse because JS has {hello: 'world'}. In fact that's an inconsistency of JS that only holds for _some_ keys (basically, valid identifiers): if the key has a space or a dash or something, you need to quote it: {'hello-1': 'world'}. So Python is more consistent in having only one style of map literal. >in python you also can't do json.hello like in js, it would be json['hello'] Again, irrelevant. Python has a different syntax. Python also has tons of stuff you can't do in JS. Operator overloading for example, in python myObj + myOtherObj is a valid user defined operation. None of this means JS or Python is superior. Ans none of this has anything to do with JSON. If you mean that JSON seems to be a better fit for JS syntax, that's because it was designed to be closer to JS syntax. That said, it's not JS syntax. {foo: "bar"} is valid JS but not JSON. JSON also doesn't accept classes, closures, and tons of other things that can go in a JS object.
- juliusmusseau 7y agoBecause of this article (which I encountered a year ago) I would say Parsing JSON is no longer a minefield. I had to write my own JSON parser/formatter a year ago (to support Java 1.2 - don't ask) and this article and its supporting github repo (https://github.com/nst/JSONTestSuite https://github.com/nst/JSONTestSuite) was an unexpected gift from the heavens.
- AgentOrange1234 7y agoWait. How is this no longer a minefield just because there is a test suite that identifies some tricky cases? Doesn’t the test suite’s matrix demonstrate that there are tons of cases that aren’t handled consistently across these parsers?
- Dylan16807 7y agoIt's a very clear list of mistakes to avoid and areas where you can choose to be lenient in parsing. If you're emitting JSON you can skim the list and avoid all of them. Either way the minefield proper is no longer your problem.
- juliusmusseau 7y agoGood point. I am presuming the test suite is comprehensive. Does it cover 100% of all JSON mines? Probably not. But it surfaced about 30 bugs in my own implementation - things I would have never dreamed of. So it certainly helped me. And just based on how thorough and insane the test suite is, I think I'm in good hands. Not perfect hands - but definitely a million times better than anything I would have come up with on my own. The test suite made my parser blow up many times, and for each blow up I got to make a conscious decision in my bugfix: how do I want to handle this? (I decided to let the 10,000 depth nested {{{{{{{{{{{{{{{{{{"key","value"}}}}}}}}}}}}}}}}}}} guy blow up even though it is legal. Yes, I'm too lazy to implement my own stack.) :-)
- deleted 7y ago[deleted]
- iamleppert 7y agoCheck out the simd JSON project if you’re interested in a super fast JSON parser: https://github.com/lemire/simdjson https://github.com/lemire/simdjson I’ve been using to process and maintain giant JSON structures and it’s faster than any other parser I’ve tried. I was able to replace my previous batch job with this as it gives real-time performance.
- calcifer 7y agoThis seems to have nothing to do with the article though?
- iamleppert 7y agoIt’s a JSON parser?
- kthejoker2 7y agoHow does it do on the article's test suite?
- iamleppert 7y agoI haven’t tested it but it parses all my JSON just fine
- glangdale 7y ago[ Original designer of much of simdjson here ] We haven't used that particular suite, but almost everything in that suite is something we've thought about. In many cases we do the right thing by not innovating and randomly allowing stuff that isn't in the spec. I see exactly one thing we didn't think about, as our construction of a parse tree is pretty basic and we don't build an associative structure even when building up an object - thus we would not register an error when confronted with the malformed input listed under "2.4 Objects Duplicated Keys", but happily build a parse tree with duplicated keys (which will be built up strictly as a linear structure, not an associative one). There seems to be leeway on this point as to what an implementation should do. It certainly doesn't fit our usage model very well to build a associative structure right there on the spot - some of our users wouldn't want that much complexity/overhead.
- Multicomp 7y agoThis might be throwing a lit match into a gasoline refinery, but why not opt for XML in some circumstances? Between its strong schema and wsdl support for internet standards like soap web services, XML covers a lot of ground that Json encoding doesn't necessarily have without add-ons. I say this knowing this is an unfashionable opinion and XML has its own weaknesses, but in the spirit of using web standards and LoC approved "archivable formats", IMO there is still a place for XML in many serialization strategies around the computing landscape. Json is perfect for serializing between client and server operations or in progressive web apps running in JavaScript. It is quite serviceable in other places as well such as microservice REST APIs, but in other areas of the landscape like middleware, database record excerpts, desktop settings, data transfer files, Json is not much better or sometimes even slightly worse than XML.
- deleted 7y ago[deleted]
- hombre_fatal 7y agoNot the best context to suggest XML superiority: https://cheatsheetseries.owasp.org/cheatsheets/XML_Security_Cheat_Sheet.html https://cheatsheetseries.owasp.org/cheatsheets/XML_Security_... If parsing JSON is bad, XML is a clusterfuck.
- nullwasamistake 7y agoJSON sucks. Maybe half our REST bugs are directly related to JSON parsing. Is that a long or an int? Boolean or the string "true"? Does my library include undefined properties in the JSON? How should I encode and decode this binary blob? We tried using OpenApi specs on the server and generators to build the clients. In general, the generators are buggy as hell. We eventually gave up as about 1/4 of the endpoints generated directly from our server code didn't work. One look at a spec doc will tell you the complexity is just too high. We are moving to gRPC. It just works, and takes all the fiddling out of HTTP. It saves us from dev slap fights over stupid cruft like whether an endpoint should be PUT or POST. And saves us a massive amount of time making all those decisions.
- hu3 7y agoOff-topic but I'd want to work on a place where half the REST bugs are from JSON parsing.
- nullwasamistake 7y agoJust get a boring webapp job in CRUD world :)
- craigds 7y agoYeah I don't believe I've ever seen a json parsing problem in 11 years of software development.
- chairmanwow 7y agoI have had the absolute joy of working with gRPC services recently. Static schemas and built in streaming mechanics are fantastic. It definitely removes a lot of my gripes with REST endpoints by design.
- userbinator 7y agoI suppose you could saw that parsing any text-based protocol in general "Is a Minefield". They look so simple and "readable", which is why they're appealing initially, but parsing text always involves lots of corner-cases and I've always thought it a huge waste of resources to use text-based protocols for data that's not actually meant for human consumption the vast majority of the time. Consider something as simple as parsing an integer in a text-based format; there may be whitespace to skip, an optional sign character, and then a loop to accumulate digits and convert them (itself a subtraction, multiply, and add), and there's still the questions of all the invalid cases and what they should do. In contrast, in a binary format, all that's required is to read the data, and the most complex thing which might be required is endianness conversion. Length-prefixed binary formats are almost trivial to parse, on par with reading a field from a struture.
- juliusmusseau 7y agoConsider only this: "1.001" I'll use JavaScript numeric literals here as my translation medium (ironic!): Norway locale parses it to: 1001 USA locale parses it to: 1.001 France locale parses it to: NaN https://docs.oracle.com/cd/E19455-01/806-0169/overview-9/index.html https://docs.oracle.com/cd/E19455-01/806-0169/overview-9/ind...
- mort96 7y agoNo. In a programming context, any norwegian or french programmer would expect that to evaluate to 1.001, not 1001 or NaN.
- eitland 7y agoSweet summer child, isn't that what people say these days? I and you can agree that the reasonable thing to do is to accept American encoding. (The below is somewhat simplified, the way I remember it on a late Saturday night 10+ years later.) The outsourced team of programmers from our software vendor did not. They (mostly) used the built in regional settings in Windows (best practice, don't reinvent the wheel) meaning we had to come up with ways to make sure the machines ran with wrong regional settings (since a lot of stuff was already serialized that way && critical parts of the software was hardcoded to use US standard.) Fun times ;-)
- carapace 7y agoASN.1 > Abstract Syntax Notation One (ASN.1) is a standard interface description language for defining data structures that can be serialized and deserialized in a cross-platform way. It is broadly used in telecommunications and computer networking, and especially in cryptography. https://en.wikipedia.org/wiki/Abstract_Syntax_Notation_One https://en.wikipedia.org/wiki/Abstract_Syntax_Notation_One Or keep re-inventing the wheel. It's not like the people paying you will notice or care, eh?
- nimish 7y agoAsn.1 is incredibly hard to actually implement . There are dozens of cases of security bugs based on bad parsers. Also there are a dozen different encodings of asn.1 data including json (JER). Its age also means that it has a bunch of obsolete datatypes. Protobuf and friends have most of the power without a lot of the drawbacks.
- carapace 7y ago> Asn.1 is incredibly hard to actually implement. For whom? > There are dozens of cases of security bugs based on bad parsers. You're saying this on a thread called "Parsing JSON Is a Minefield", eh? In any event, this is not unique to ASN.1. I haven't checked but I don't doubt there are similar cases for Protobuf, etc. > Also there are a dozen different encodings of asn.1 data including json (JER). So what? That's the opposite of a problem. > Its age also means that it has a bunch of obsolete datatypes. So don't use them. - - - - My point is that if the time and effort that was spent on Protobuf and CapnProto and all the others had somehow been spent instead on perfecting ASN.1 then, uh, that would have been good...
- kentonv 7y ago> My point is that if the time and effort that was spent on Protobuf and CapnProto and all the others had somehow been spent instead on perfecting ASN.1 then, uh, that would have been good... I wrote proto2 in 20% time at Google and I developed Cap'n Proto entirely on my own time, unpaid. If you think ASN.1 could be perfected with a similar amount of work then why don't you do it?
- ufo 7y agoThe crashing test cases look scary from a security perspective, specially in the C-based parsers. Does anyone know if these results are still up to date or if the bugs have already been fixed?
- SigmundA 7y agoBeen of the opinion a while a lot of issues could be resolved if we agreed on a streamable binary format that had good definitions for data types (including integers and dates). String formats are great an all for viewing in whatever text viewer but so inefficient and then you have the whole escaping string inside of strings and string encoding binary data. If we all agreed on a binary format then there would be a viewer for it in every debugging tool. ASN.1, Protobuf, BSON, ION. MSGPACK whatever. I would prefer a binary format that doesn't repeat keys for efficiency where the schema can be sent separately or inlined. But even one that's basically binary JSON with more types would be step up.
- failrate 7y agoParsing is a minefield. General purpose computing systems are minefields. Of all human readable formats I've ever worked with, only S-expressions have proven easier and safer to parse. Json.org even has unambiguous railway diagrams!
- filoeleven 7y agoI wish EDN would catch on. The simplicity of JSON with better number handling, a few more very useful data types like namespaced keywords and sets, arbitrarily complex keys, and an even terser, more readable syntax. [k v k v] beats [k: v, k: v] hands down, and you can use whitespace commas if you want them. https://github.com/edn-format/edn https://github.com/edn-format/edn
- rendaw 7y agoPlugging my amazing JSON-like format! https://gitlab.com/rendaw/luxem https://gitlab.com/rendaw/luxem In case anyone doesn't click the link, its much simpler than JSON, leaves interpretation up to the reader, supports polymorphic data, and has some tweaks to make it nicer to edit by hand. I've used this in a bunch of personal projects and maps all the models I've come across perfectly. If you need more power in your format you're better off using Lua than YAML. I have a Rust Serde implementation 90% complete I could finish up if anyone wants it.
- mkl 7y agoTrailing commas are of course very sensible! This doesn't seem simpler than JSON otherwise, though, e.g. type declarations and optional quotes. Why is leaving interpretation up to the reader desirable? Shouldn't things always come out the same? Asterisks are an unusual choice of comment syntax. What if your comment needs to contain an asterisk? Why not "//..." and/or "/* ... */", or "#..."?
- ape4 7y agoRelaxed JSON is pretty good. http://www.relaxedjson.org/ http://www.relaxedjson.org/
- krispbyte 7y agoRelax it a bit more and you get Neon: https://ne-on.org/ https://ne-on.org/
- inopinatus 7y agoOnce you're parsed the first minefield, another crop emerges: interpreting the result. Even the range of values seen in the wild for a supposedly simple boolean attribute is just mind-boggling. Setting aside all the noise from jokers trying it on with fuzzing engines, we'll see all of these presented to various APIs: true false null 0 | 1 "true" | "false" (with assorted variation by "yes" | "no" case and initial character) "" | "0" | "1" "\u2713" (hi DHH) -1 (with complements) "[object Object]" { "value": true } (and friends) (attribute not present) "敵牴" That last looks like a doozy, but old lags will guess what's going on right away. It's the octets of the 8-bit string "true", misinterpreted as UCS-2 (16-bit wide character) code points and then spat out as UTF8. Google translates it, quite appropriately, as "Enemy". Oddly though, according to my records, never seen a "NULL".
- rurban 7y agoAgain. Parsing is a minefield in general, but parsing JSON is one of the easiest tasks of all serialization formats. It's also the only secure format. Its various spec bugs (by omission) are not that dramatic, and the various "enhancements" only made it worse, ie more insecure. Still, bad but not a minefield. What worries me most that my JSON module is the defacto perl standard, passes all these tests, was the very first to add all these tests, is the fastest, and still is not included in that list, just some outdated modules which should not be used at all. Checking best practices besides maintaining a spec obviously also is a minefield.
- rurban 7y agoI forgot another major JSON minefield problem which is not mentioned nor tested here: stackoverflow. This is in fact the most important problem to test against, because it might lead to exploitable stack ROP. JSON is usually parsed recursively, and deeply nested structures are mostly not depth counted. One can trivially construct a nested array or map of 500 to 30000 elements, and at one point the parser either fails or crashes with an overflow. This number is fixed, thus trivially exploitable. The test spec should contain the max. depth for arrays and maps, and if there's a fixed builtin limit, a compile-time limit, implicit limit by crash, or none. non recursive parsers are fine.
- cockatoo2 7y agoAnimals trampling snow? That's some horse shit
- jstewartmobile 7y agoThis is all well and good, but which decoders are already in the browser? XML and JSON last I checked.
- peterwwillis 7y agoThe product owner perspective on this should be "nothing supports anything unless it is tested". If you pick up a standard and just assume other products will be able to work with it, you're in for a surprise. I don't care if it's TCP sockets or .ini files; if you didn't test compatibility with the product you expect will interact with yours through the standard, consider it unstable, and don't advertise support for it. Sometimes you have to support a standard itself, like WPA2, so you implement the standard according to internet engineering best practice: be liberal with what you accept, and conservative with what you transmit (or something to that effect). Then test compatibility with the major products you know will want to use it, and fix the bugs you find.
- jasonhansel 7y agoCompared to XML, Markdown, and other human readable formats, this is...actually not too bad. I was expecting worse.
- mirimir 7y agoI'm not a professional coder. And I mostly work with tabular data, in spreadsheets and SQL. I like to get my data as delimited text files. Ideally, delimited with some character that's 100% guaranteed to never occur in the data. In my experience, "|" is often a good option, but you never know. And CSV, even with quotes, can be a nightmare, especially if the data contains addresses. Or names with quoted nicknames. Anyway, given the choice, I always pick JSON over XML. Because with JSON, I can always identify the data blocks that I need, and parse them out with bash and spreadsheets. Not with XML, however. Just as not with HTML.
- warmfuzzykitten 7y agoI certainly would not criticize such a thorough examination for being facile, but I do want to point out that the conclusion "But sometimes, simple specifications just mean hidden complexity" is not supported by the article. Almost all of the end cases are caused by implementors ignoring or extending the simple specification.