20 ms·
Parsing JSON is a Minefield
- kowdermeister 10y agoI still love JSON regardless :) Client / server side languages have first class support for serialization and in most cases the data structures are rather easy. I'd be very skeptical if one would suggest an alternative format for a web based project, however I can imagine such situations.
- michaelp983 10y agoIf you could cut your server spend to ~1/3 - 1/10th (depending on the language and application complexity) I think it would be worth it. All depends on your needs. If I am building a dummy app, yes I'll use JSON. If I am building a real application, I am building it with FlatBuffers or SBE.
- mi100hael 10y ago> In conclusion, JSON is not a data format you can rely on blindly. That was definitely not my take-away from the article. More like "JSON is not a data format you can rely on blindly if you are using an esoteric edge-case and/or an alpha-stage parsing library." I haven't ever run into a single JSON issue that wasn't due to my own fat fingers or trying to serialize data that would have been better suited to something like BSON.
- Tepix 10y agoYou are not considering a hostile environment like the internet where an attacker will use edge cases to get unforeseen results.
- inimino 10y agoThe takeaway is "JSON parsing can fail and possibly crash" but I don't think that should be news.
- mi100hael 10y agoI couldn't really care less if someone POSTs some garbage JSON that results in them getting a 500 response. Better than someone POSTing an XML bomb and affecting other peoples' requests. Please enlighten me if you know of a serialization format with libraries for all common languages that lacks any gotchas or edge cases.
- firethief 10y agoIf all you're doing is parsing their JSON for it's own sake, it's just a 500; but that's the boring case. Consider what happens when typical web code is interacting with the JSON parser.
- Someone 10y agoAll software has edges, so edge cases are unavoidable. The best you can do is: - interpret the spec to the letter. - for every fragment of a statement you write, consider whether it might conceivably go wrong, and handle those cases (in the simplest matter because 'handling' means writing code, and that code, too, needs to go through this process). For example, a json parser must be prepared to handle missing values, extremely long keys and values (integers may have thousands of digits, think long about the question whether 64 bits always is enough for storing a string length, etc.), etc. - if you are truly paranoid, have very stringent security requirements, or expect to be heavily attacked, run the parser in a separate process. - fuzz your implementation.
- gmfawcett 10y ago> think long about the question whether 64 bits always is enough for storing a string length, etc.), etc. I'm struggling to think of any realistic scenario where this isn't true!
- Someone 10y agoSo do I, but you should still consciously decide whether to add an overflow check. Let's do a quick estimate: one can read in the order of 2^24 bytes/second from disk. A day has on the order of 2^16 seconds, so that's 2^40 bytes/day. => You will need 2^24 days to read 2^64 bytes. I think that's around 50k years. That an attacker will try to generate a buffer overflow this way is a risk I would take, even if I thought the hardware had room to store that string. The only way I can foresee a real risk is when an optimizer can optimize away the computation of a string whose length it is asked to compute by an attacker. That's still very much far-fetched, and if no string gets allocated it's hard to see how it could become a security issue, but it could be a reason to be extra careful, for example when providing online access to a C++ compiler, with its template metaprogramming capabilities.
- empath75 10y agoYeah, basically the json parsers and generators and the programs that use them seem to have settled on implementations that mostly work together.
- jerf 10y agoThe sum totality of all the issues raised in that post is not an esoteric edge case, even if each individual element is an esoteric edge case. If you haven't encountered any of them in your real code yet, there's two basic possibilities. Either you aren't using JSON very hard at all... or you have encountered them and you just didn't realize it. You will, sooner or later. I'm not saying JSON is bad. Personally I think the sloppiness that this post enumerates is part of its success, and part of why it has so much support in so many languages. The more you nail down the semantics, the harder such widespread support gets. Right now you get a nice subset of the JSON spec that works in a huge variety of languages, where the JSON support for a given language usually converts the JSON into something fairly native for the language. This wouldn't be possible if you nailed down the semantics as hard as something like Protobuf does. For instance, a ton of statically-typed languages will let you encode and decode integers into JSON. But JSON doesn't have an integer type. Strict JSON support ought to forbid integers. But it's sloppiness lets us all just sort of ignore that and get on with life. If you have a need for precision, avoid JSON. And people probably generally need more precision than they realize and probably ought to reach for JSON a bit less often than they do. But on the other hand, the whole thing does mostly work, right? That can't be ignored.
- inimino 10y agoJSON doesn't have an integer type, but it certainly supports integers. Within, obviously, implementation defined limits. I'm with you up to "if you need precision, avoid JSON". Actually JSON is fine for the kinds of precision most use cases require, and when it isn't, you probably know it.
- lobster_johnson 10y agoJSON's number support is a source of problems, because the standard is so informal [1]: JSON is agnostic about numbers. ... JSON instead offers only the representation of numbers that humans use: a sequence of digits This poses a problem for some languages, and tends to break things. You would expect encode(decode(string)) == string, but languages deal with numbers differently. For example, in Go, if you decode into a map[string]interface{}, you will get float64 by default. If you decode "42" and then encode it back to JSON, you'll get "42.0". (This is the reason Go has a special type you can use, json.Number, that preserves the value as its original string value.) This mostly causes issues for code that needs to be data-structure-agnostic (in the Go case: decoding into a interface{}, rather than a struct with "json:" tags), but there are edges cases where you can get a surprise. For example, since all numbers are technically both integers and floats, something like {"visitorCount": 42.0} is perfectly valid, and a client has to know to coerce the number into an integer if it wants to deal with it sanely, even though the meaning of that number might be nonsensical if treated as a float. [1] http://www.ecma-international.org/publications/files/ECMA-ST/ECMA-404.pdf http://www.ecma-international.org/publications/files/ECMA-ST...
- wtetzner 10y ago> In conclusion, JSON is not a data format you can rely on blindly. Seems like an odd thing to say in general. Perhaps "JSON parsers are not libraries you can rely on blindly", but that seems true of everything.
- romaniv 10y agoIf there are two servers/services that communicate via JSON and they use different parsers, these types of issues can lead to rather nasty problems. Even if both parties fail gracefully on their own. This gets even worse if your software is an integration layer between two services you do not control.
- indexerror 10y ago> In conclusion, JSON is not a data format you can rely on blindly. What does HN suggest for configuration files (to be written by a human essentially)? I am looking at YAML and TOML. My experience with JSON based config files was horrible.
- FreeFull 10y agoIn my personal experience TOML works really well. It's a little reminiscent of .ini files, but definitely is better.
- jzwinck 10y agoYAML or TSV depending on whether your configuration looks like a rectangular table. If you want extreme flexibility using C++ as the main language, take a look at my project: https://github.com/jzwinck/pccl https://github.com/jzwinck/pccl It lets you configure your C++ apps using Python. Config items can even be Python functions.
- kozhevnikov 10y agohttps://github.com/typesafehub/config/blob/master/HOCON.md https://github.com/typesafehub/config/blob/master/HOCON.md
- marios 10y agoI don't have a specific recommendation, but when I see a project uses a JSON file as configuration, I wonder: "hasn't the author ever needed to include a comment in the configuration ?".
- mmagin 10y ago"NaN and Infinity" Yeah. And I learned this the hard way with the Perl module JSON::XS. It successfully encodes a Perl NaN, but its decoder will choke on that JSON. (Reported it to the maintainer who insists that is consistent with the documentation and wouldn't fix it)
- TazeTSchnitzel 10y agoSimilarly, Python's encoder violates the JSON specification by default, as it produces `Infinity`, `NaN` and `-NaN`, which other JSON parsers choke on.
- joesb 10y agoI don't get it. Why would unaware JSON parsers choke on `Infinitiy`, `NaN` or `-Nan`? JSON has no concept of schema. So if a parser sees "Inifinity", which it doesn't have any concept of, why would it do anything except treating that as string of a word "Infinity"?
- cdmckay 10y agoBecause it's not quoted like a string, it's a literal
- rurban 10y agoNobody should use JSON::XS anymore, everybody switched to Cancel::JSON::XS which does have those features, and less bugs. I just added those great testcases, and found previously unknown bugs. But the python test runner from this repo gives a few false negatives, PR coming soon. It was much easier to write it in perl.
- benibela 10y agoOh no. My parser/serializer does the same. How could that be fixed? Anything better to serialize NaN to than NaN? Or parse NaN?
- jayd16 10y agotl;dr JSON with a bunch of shitty extensions is awful. The error handling among JSON parsers is inconsistent.
- s_q_b 10y agoWell, first and most obviously, if you are thinking of rolling your own JSON parser, stop and seek medical attention. Secondly, assume that parsing your input will crash, so catch the error and have your application fail gracefully. This is the number one security issue I encounter in "security audited" PHP. (The second being the "==" vs. "===" debacle that is PHP comparison.) As one example, consider what happens when the code opens a session, sets the session username, then parses some input JSON before the password is evaluated. Crashing the script at the json_decode() fails with the session open, so the attacker can log in as anyone. Third, parsing everything is a minefield, including HTML. We as a community invest a lot of collective effort in improving those parsers, but this article does serve as a useful reminder of a lot of the infrastructure we take for granted. Takeaways: Don't parse JSON yourself, and don't let calls to the parsing functions fail silently.
- lucb1e 10y ago> Well, first and most obviously, if you are thinking of rolling your own JSON parser, stop and seek medical attention. Been there done that. (The medical attention, I mean.) Worked just fine. The article makes it sound extremely difficult, but 100% of the article is about edge cases that rarely happen with normal encoders and can often be ignored (e.g. who cares if you escape your tab character?). > consider what happens when the code opens a session, sets the session username, then parses some input JSON before the password is evaluated Edit: I responded more elaborately on the unlikelihood of this, but honestly, I can't come up with a single conceivable scenario. How would you decode part of the JSON and only parse the password bit later?
- wtetzner 10y ago> Setting the session flag for "this user is logged in" before checking (or even decoding!) the password seems rather backwards to me. Yeah, that seems like a problem regardless of whether or not you're parsing JSON.
- rch 10y agoProbably a symptom of the PHP multiverse: anything that can happen, has happened.
- Tepix 10y agoThe important lessen is that you can't blindly rely on your JSON parser to save your ass when you are dealing with untrusted input. If sending 1000 "["s will crash your application, you have a problem. I hope the JSON parser authors will improve their parsers.
- newsat13 10y agoAnyone else seeing forbidden? Forbidden You don't have permission to access to this document on this server.
- gt2 10y agoyes, same here
- lucb1e 10y agoYou are not the only ones, but it works for me. Google cache link is posted elsewhere in the thread: https://news.ycombinator.com/item?id=12797047 https://news.ycombinator.com/item?id=12797047
- GuzmanMan 10y agoI saw that, but when I refreshed the page I was able to load the content.
- beefburger 10y agoToo heavy CPU load. Now I turned the PHP page into a static HTML one. Configuring web servers is a minefield :)
- metafunctor 10y agoThe page has been taken down for some reason (getting a 403). Google cache: http://webcache.googleusercontent.com/search?q=cache:8jVuBmxKEkQJ:seriot.ch/parsing_json.php http://webcache.googleusercontent.com/search?q=cache:8jVuBmx...
- lucb1e 10y agoWorks for me. Edit: But you are not alone. Elsewhere in the thread: https://news.ycombinator.com/item?id=12797032 https://news.ycombinator.com/item?id=12797032
- metafunctor 10y agoYep, looks like it was down only for a few minutes.
- beefburger 10y agoToo heavy CPU load. Now I turned the PHP page into a static HTML one and it's back online. Configuring web servers is a minefield :)
- andrewvijay 10y agoBad thing to read when I'm writing a sass to json module
- lucb1e 10y agoThe encoding takeaway seems simple: escape everything with \uxxxx characters that is outside of the ASCII range /[ -~]/ (regex) and you'll be pretty much fine. Set the encoder to utf-8, don't leave [dangling,commas,], and a few other things that are obvious from json.org.
- inimino 10y agoEscaping everything outside of ASCII is terribly human-unfriendly and wasteful in bytes as well.
- nathancahill 10y agoF
- amelius 10y agoOne of the biggest flaws of JSON is that it doesn't support "undefined". This makes translating Javascript structures to and from JSON actually not preserve the original value. Sigh.
- indubitably 10y agoBut undefined is specific to Javascript… there are lots of other Javascript things that JSON doesn't handle either, like Set or Map objects. It's not intended to serialize arbibtrary JS objects — it's intended to serialize a useful least-common-denominator which has proven useful through experience.
- inimino 10y agoThere are lots of reasons beyond this one why arbitrary JS values can't roundtrip through JSON. This is a conscious design decision, not an oversight.
- ohstopitu 10y agoWhen I didn't know better, I wrote my own JSON parser for Java (it was years back and I didn't know about java libraries). From experience: DON'T. DO. IT. That said, if you have decided to do it.... 1) know fully well that it'll fail and build it with that assumption. 2) Please, please, please...give useful error messages when it does fail or you'd be spending way too much time over something simple.
- SloopJon 10y agoFigures that something like this would be posted on my day off. I put this through a parser that I cover, and found that the only failures were for top-level scalars, which we don't support, and for things we accept that we shouldn't. I'll look through the latter tomorrow, as well as the optional "i_" tests. Test suites are a huge value add for a standard, so thank you, Nicolas, for researching and creating this one. I was surprised that JSON_checker failed some of the tests. I use its test suite too.
- mjpa 10y agoWrote my own JSON parser (https://github.com/MJPA/SimpleJSON https://github.com/MJPA/SimpleJSON) a while ago... not sure how it's a minefield unless I'm missing something?
- valarauca1 10y agoDid you read the post? The JSON standard(s) are very simple. But in practice this simplicity leads to a lot of edge cases. Very deeply nested structures, numbers that run on to infinity. The point of this post is testing various parsers against these malicious structures. See this image: http://seriot.ch/json/pruned_results.png http://seriot.ch/json/pruned_results.png
- mjpa 10y ago"In conclusion, JSON is not a data format you can rely on blindly." - that suggests the format is bad when it's highlighting problems with the parsers. I see it like saying "Plain text is a bad format because notepad bails on large files"
- valarauca1 10y agoIn conclusion, NOUN1 is not a NOUN2 you can rely on blindly This is true for everything. I don't see why this is evidence of anything. If you rely on any system blindly you are doing something wrong.
- SloopJon 10y agoPerhaps you'd be interested to know that your JSONDemo program fails the following tests: hang: y_number_huge_exp.json segfault: n_structure_100000_opening_arrays.json n_structure_open_array_object.json fail: n_number_then_00.json n_string_unescaped_tab.json n_structure_capitalized_True.json
- mjpa 10y agoI would :) Was going to run the tests myself later but needed a box with python3 on it!
- Confusion 10y agoThere was a great article at some point that explained why 'be liberal in what you accept' is a very bad engineering practice in certain circumstances, such as setting a standard, because it causes users to be confused and annoyed when a value accepted by system A is subsequently not accepted by supposedly compatible system B. Leading to pointless discussions about what the spec 'intended' and subtle incompatibility. Anyone know what article I mean?
- webmaven 10y agoPerhaps "The Harmful Consequences of Postel's Maxim"?: https://news.ycombinator.com/item?id=9824638 https://news.ycombinator.com/item?id=9824638 If not, perhaps one of these: http://programmingisterrible.com/post/42215715657/postels-principle-is-a-bad-idea http://programmingisterrible.com/post/42215715657/postels-pr... https://bitworking.org/news/There_are_no_exceptions_to_Postel_s_Law_ https://bitworking.org/news/There_are_no_exceptions_to_Poste... http://trevorjim.com/postels-law-is-not-for-you/ http://trevorjim.com/postels-law-is-not-for-you/
- Confusion 10y agoYes, thank you, it was the one by 'programmingisterrible' and the linked paper by Patterson, Sassaman, and Bratus.
- inimino 10y agoI've seen this argument most frequently made with regards to XML.
- collyw 10y agoThats pretty much my experience when building software as well. A lot of the time I have been liberal to incorporate legacy data, and every time in has ended up being the cause of the majority of bugs in the systems I have built.
- peatmoss 10y agoWhat ever happened with EDN (pronounced "eden") from the Clojure people? https://clojure.github.io/clojure/clojure.edn-api.html https://clojure.github.io/clojure/clojure.edn-api.html https://github.com/edn-format/edn https://github.com/edn-format/edn I always thought that seemed like a nice alternative data format to JSON. Anyone using this it in the wild?
- arohner 10y agoClojure programmers use it everywhere. I suspect almost nobody else does though.
- tragic 10y agoEven if you're a clojure shop, you've got the issue that at your system boundary, everything else the world accepts json and/or XML, and nothing supports EDN/transit. So your data needs to be serialisable to one of those anyway, at least if it crosses the boundary. There is such a thing as a network effect, even in data serialisation formats...
- grayrest 10y agoIt got supplanted by Transit [0] as an interchange format. Both are only really used within the Clojure community. I use Transit for internal-facing APIs. [0] https://github.com/cognitect/transit-format https://github.com/cognitect/transit-format
- dep_b 10y agoNow the mess that is called JavaScript dates has crept into any system imaginable in the world. I can understand we needed to go for the lowest denominator but Crockford's card really could cram in another line with a date time string format.
- petre 10y agoJust use epoch or ISO1601.
- dep_b 10y agoYou're preaching to the choir here. Hell is other people's code.
- inimino 10y agoI think adding a date type would have probably doubled the complexity of JSON and taken it right out of the sweet spot that has made it popular.
- niftich 10y agoThis. Datetime is hard, because RFC 3339 is underspecified [1], and TOML wrestled with this a lot. [1] https://news.ycombinator.com/item?id=12364393#12364805 https://news.ycombinator.com/item?id=12364393#12364805
- halomru 10y agoI thought strings with ISO8601-encoded dates was the de-facto norm (i.e. "20161026T2044+0200 for this comment's timestamp). Have I been living in a bubble with the majority being less sane?
- metaloha 10y agoAm I wrong in seeing that PHP seems to fail the least weirdly in the full results?
- TazeTSchnitzel 10y agoIt's more complicated in practice. Prior to PHP 7.0, the official JSON extension was replaced with a completely different one in many distributions due to licensing issues, and since PHP 7.0, PHP's official JSON extension is yet another completely different implementation.
- DanielRibeiro 10y agoWow! This was a great practical analysis of existing implementations, besides a great technical overview of the spec(s). Thanks for open sourcing the analysis code[1], and for the extended results[2] [1] https://github.com/nst/JSONTestSuite https://github.com/nst/JSONTestSuite [2] http://seriot.ch/json/parsing.html http://seriot.ch/json/parsing.html
- edem 10y agoJSON is the de facto standard when it comes to (un)serialising and exchanging data in web and mobile programming. I disagree. Take protobuf for example. You get schemas, data structures, and a parser in one package which is actually a lot smaller than JSON and compiles to nearly all the commonly used languages. Ever since I've started using it my life became so easier! If you don't want your data to be human readable (which is very common) you should not use JSON as a data interchange format.
- djur 10y agoJSON is still the de facto standard, regardless of whether it should be.
- iamatworknow 10y agoPlease tell that to the multi-billion dollar company I'm currently working to integrate with that decided 2016 was a good year to implement a brand new SOAP API.
- edem 10y agoI don't think so. Can you back it up with facts?
- oldmanjay 10y agoSince you haven't demonstrated your premise beyond talking about what you personally prefer, that particular ball is still in your court. The world isn't obliged to accept your unsupported opinions as truth until you're convinced otherwise.
- edem 10y agoI don't need to demonstrate anything. It is not de facto standard since if it were everybody were using it which is not the case (Google is the best example). Look up what the term means and you will understand.
- cbhl 10y agoIn practice, people use increasingly smaller subsets of JavaScript to transmit data. For example, a common pattern is to transmit (numeric) user IDs as strings so that they don't get mangled by floating-point precision issues with large numbers. You see both Twitter and Facebook APIs do this, for example.
- Senji 10y agoModern languages should feature a GUID type in the standard library.
- zodiac 10y agoI've read that you should treat IDs as strings anyway because you want to discourage incrementing IDs, adding IDs together etc. Anything else you want to do with integer IDs, e.g. comparing, can be done with sting IDs as well.
- novaleaf 10y agowhen parsing human constructed JSON, use JSON5 for the win. http://json5.org/ http://json5.org/
- kstenerud 10y agoI did write my own parser, but for a reason: I need it to be able to recover as much data as possible from a damaged, malformed, or incomplete file. Turns out that a good chunk of these tests are for somewhat malformed, but not impossible to reason about files. Extra commas, unescaped characters, leading zeroes... I'd rather just accept those kinds of things rather than throw an error in the user's face. It's a big bad world out there, and data is by definition corrupt. And this is borne out when I plug my parser into this test suite: Many, many yellow results, which is exactly how I want it.
- redleggedfrog 10y agoCrockford needs to write "JSON, The Good Parts."
- seagreen 10y agoNo he doesn't. "JSON, the good parts" is just JSON. The problems the post is describing have to do with RFC 7159 allowed extensions.
- realkitkat 10y agoIf JSON is comparable to minefield, then I guess XML and ASN.1 are nothing short of nuclear Armageddon in complexity and ones ability to shoot themselves into the leg ;-)
- Impossible 10y agoAlso, as someone who has written an XML parser, according to some of the comments in this threads I'm way beyond medical help, and should give up on life :).
- lameexcuse333 10y agoDepends on when you did this crazy thing. Was it in the dark days before they were "ubiquitous"? (for lack of a better word. also it's just a fun word) Then you're a hero to those who made use of it. If someone were to do that today, though, then yes, seek help. ;)
- ocschwar 10y agoWell, every person you spared from that experience owes you a beer...
- twic 10y agoI got dinged in a thread the other day for writing my own JSON parser, so i'm not sure i should confess to also having written enough of an ASN.1 parser to deal with certain PEM key formats: https://github.com/pivotal/cf-env/blob/master/src/main/java/io/pivotal/labs/cfenv/crypto/KeyAlgorithm.java https://github.com/pivotal/cf-env/blob/master/src/main/java/...
- deleted 10y ago[deleted]
- xamuel 10y agoHave to take HN anti-parser comments with a grain of salt. I wrote a parser for a format that everyone here would crucify me for. When I went on the job market, I removed it from my github, fearing employers would see it as a black mark. Soon after I got email from the creator of a YUUUGE project (as in, 5 digits of stars on github) asking why I had removed this parser that they were using ...
- vhost- 10y agoYou definitely can't rely on it. Just the other day I was given a task to take a request payload from our front end and do some stuff with it on our backend. The payload looked like this: {"thing": [{"values": ["foobar"], "type": "blah blah"}, "some identifier"], "other thing": "some string"}. It's mixing types in arrays which is problematic for most statically types languages. Tips for Go: Don't use map[string]interface{} and circumvent the type system (I've seen this a lot in production). The fix involves the UnmarshalJSON and MarshalJSON interfaces. This lets you put the data into a structure that's sane and re-encode it back to something the other system expects.
- austincheney 10y agoWriting parsers is hard and takes some experience, but its not as hard or as impossible as most of these comments make out. JSON is retarted simple to parse, even in the face of certain edge case ambiguities. I can say this from experience after having written an HTML/XML parser that provides support for various template schemes: Twig, Elm, Handlebars, ERB, Apache Velocity, JSP, Freemarker, and many more. I have written a JavaScript parser that supports React JSX, JSON, TypeScript, C#, Java, and many more things. In years I have been programming I frequently hear whining like, "its too hard". Don't care. While you are wasting oxygen crying about how hard life is somebody else will roll a solution you will ultimately consume.
- aikah 10y agowhere are all these goodies you boast about then ?
- austincheney 10y agohttp://prettydiff.com/ http://prettydiff.com/
- paulddraper 10y agoLots of issues are trivially answered. --- > Scalars..In practice, many popular parsers do still implement RFC 4627 and won't parse lonely values. Right. RFC 7159 expanded the definition of a JSON text. > A JSON text is a serialized value. Note that certain previous specifications of JSON constrained a JSON text to be an object or an array. If RFC 7159 wasn't different from 4627, there'd be no reason for 7159. Same with RFC 1945 and 7230 for HTTP. (Of course, HTTP is versioned...maybe he just means to repeat the earlier versioning criticism.) --- > it is unclear to me whether parsers are allowed to raise errors when they meet extreme values such 1e9999 or 0.0000000000000000000000000000001 And then quotes the relevant part of the RFC 7159 grammar with answers the question: > This specification allows implementations to set limits on the range and precision of numbers accepted. Since software that implements IEEE 754-2008 binary64 (double precision) numbers [IEEE754] is generally available and widely used, good interoperability can be achieved by implementations that expect no more precision or range than these provide, in the sense that implementations will approximate JSON numbers within the expected precision. A JSON number such as 1E400 or 3.141592653589793238462643383279 may indicate potential interoperability problems, since it suggests that the software that created it expects receiving software to have greater capabilities for numeric magnitude and precision than is widely available. Parsers may limit this however they like. And so may serializers. This includes yielding errors. (Though approximating the nearest possible 64-bit double is IMO the better choice.) --- So yeah, in the end there is fair amount of flexibility in standard JSON. To summarize: > An implementation may set limits on the size of texts that it accepts. > An implementation may set limits on the maximum depth of nesting. [this one was never mentioned though] > An implementation may set limits on the range and precision of numbers. > An implementation may set limits on the length and character contents of strings. Most implementations on 32-bit platforms will not parse 5GB JSON texts.
- tofupup 10y agoit is an improvement but when i am in these situations i usually grab the first library out there.
- eridius 10y agoSpeaking as someone who wrote a JSON parser, this article and the accompanying test suite looks to be very valuable, and I will be adding this test suite to my parser's tests shortly. That said, since my parser is a pure-Swift parser, I'm kind of bummed that the author didn't include it already, but instead chose to include an apparently buggy parser by Big Nerd Ranch instead. My parser is https://github.com/postmates/PMJSON https://github.com/postmates/PMJSON
- ninjakeyboard 10y agounless for fun, rolling your own json parser is like writing bubble sort for use in your prod app.
- youdontknowtho 10y agoYou know what is more like a minefield...a minefield... http://www.afghan-network.net/Landmines/ http://www.afghan-network.net/Landmines/ Not trying to be a dork, but thought this would be a good place to bring up...if anyone is interested...in the usage of landmines in current conflicts and the way that they tend to linger. Call this a comment factoid. Off topic, but interesting.
- michaelp983 10y agoThere are alternatives to JSON that are OS, available in many languages, actively supported, and both CPU and memory efficient. FlatBuffers & Netflix https://www.youtube.com/watch?v=K_AfmRc-TLE&feature=youtu.be&t=21m30s https://www.youtube.com/watch?v=K_AfmRc-TLE&feature=youtu.be...
- singularity2001 10y agoJSON = require('json5') And you can even use comments!! (no comment) // http://json5.org/ http://json5.org/
- RangerScience 10y agoThis is fantastic. However, it looks like the detailed conclusion is "exactly matching the RFC is a minefield". About a month ago (for the third time, since I don't own the first two implementations) I made a very forgiving (and very error-unprotected) JSON parser: https://github.com/narfanator/maptionary https://github.com/narfanator/maptionary The core of JSON parsing, from that experience, seems really simple; it's catching all the edge cases that's hard. In any event, I look forward to taking the time to test against this test suite!
- RangerScience 10y agoDo you have an explanation anywhere of why each (or any) of the edge cases is supposed to succeed or fail, or why it commonly does what it's not supposed to do? I realize that's almost as much work as writing each test case in the first place, but even a subset of the test cases having that explanation would be valuable.
- nitwit005 10y agoPeople tend to screw up the unicode aspects more than the general parsing. And, indeed, the example JSON parser provided checks for a UTF-8 byte order mark, but doesn't validate that the data is valid UTF-8, so it will let through strings that might cause an application problems. Although there is a commented out method to validate a code point, so I guess he understood that it was an issue.
- gcirino42 10y agoThe correct answer to parsing JSON is... don't. We experimented last hackday with building Netflix on TVs without using JSON serialization (Netflix is very heavy on JSON payloads) by packing the bytes by hand to get a sense of how much the "easy to read" abstraction was costing us, and the results were staggering. On low end hardware, performance was visibly better, and data access was lightening fast. Michael Paulson, a member of the team, just gave a talk about how to use flatbuffers to accomplish the same sort of thing ("JSOFF: A World Without JSON"), linked in this thread: https://news.ycombinator.com/item?id=12799904 https://news.ycombinator.com/item?id=12799904
- dvt 10y agoNot sure what your point is (or the point of that presentation, for that matter). Of course there are binary serialization formats that are faster than XML or JSON, and of course they're less error-prone. This has been known for about 40 years now. JSON/XML are used precisely because people want a human-readable interchange format. For high-performance uses, consider Google's Protocol Buffers or Boost::serialize. You're acting like you just hackathoned the biggest thing since sliced bread, but that's exactly how payloads have been sent (until high-bandwidth made us all lazy) since the inception of the Internet.
- gcirino42 10y agoI thought my point was clear - don't get involved parsing JSON; I agree with the OP, parsing JSON is a minefield. I went further by implying that it is also unnecessary when ease of reading isn't needed, and called out some alternatives. I think it's amusing that you mentioned protocol buffers - were you aware that when I mentioned flat buffers that they were built in relation to performance inefficiencies in the very protocol buffers that you mentioned? We didn't just "hackathoned the biggest thing since sliced bread", btw, we took a real world example of exchanging a human readable format for a human-with-tools-readable one and saw a significant win. High-bandwidth also isn't as prevalent as you think, and yes, you're generally paying both performance wise and occasionally monetarily for the laziness you mentioned. But then, if you've known this for 40 years and don't know how to measure it, there's not much I can do for you in a comment.
- SFJulie 10y agoBy sheer randomness I was having the thought about it today: I made some code to highlight where the stdlib json module sees the mistakes in JSON decoding in the python stdlib. I used the exception with string "blabal at line x, col y, char(c - d)" to actually highlight (ANSI colors) WHERE the mistake were. https://gist.github.com/jul/406da833d99e545085dac2f368a3b850#file-test-png https://gist.github.com/jul/406da833d99e545085dac2f368a3b850... I played a tad with it, and the highlighted area for missing separators, unfinished string, lack of identifier were making no sense. I thought I was having a bug. Checked and re-checked. But, No. I made this tool because, whatever the linters are I was always wondering why I was not able to edit or validate json (especially one generated by tools coded by idiots) easily. I thought I was stupid thinking json were complex. Thanks.
- 77pt77 10y ago> it won't parse u-escaped invalid codepoints: ["\ud800"] How is this not expected behaviour? The string is not well-formed. Same thing with decent XML parsers. They croak when you give them invalid codepoints.
- dgreensp 10y agoAn informative article. The point is not that parsing JSON is "hard" in any sense of the word. It's that it's underspecified, which leads to parsers disagreeing. Although the syntax of JSON is simple and well-specced: * The semantics are not fully specified * There are multiple specs (which is a problem even if they are 99% equivalent) * Some of the specs are needlessly ambiguous in edge cases * Some parsers are needlessly lenient or support extensions
- chrismarlow9 10y agothis is kind of a vulnerability developers wet dream, especially that graph...
- RX14 10y agoFixes for many of the issues raised by this post have since made it into crystal master: https://github.com/crystal-lang/crystal/commit/7eb738f550818825786e90389ac84d2a2eb13e13 https://github.com/crystal-lang/crystal/commit/7eb738f550818...
- devmunchies 10y agoI love how quickly things move in crystal right now. ️
- nickpsecurity 10y agoCrap like this is why people should just use older ones that work. Some of the issues I see in the comments weren't present in Sun's XDR: https://tools.ietf.org/html/rfc4506 https://tools.ietf.org/html/rfc4506 Or even LISP s-expressions if you want organized text.
- boggydepot 10y agoSo what's the alternative? If you had a time machine and went back in time, what would you recommend/bring as an alternative to JSON?
- rahrahrah 10y agoOh boy this is another one of those threads.. Party A: X is harder than it looks Party B: X isn't as hard as you're making it look Boring.