10 ms·
A Journey building a fast JSON parser and full JSONPath, Oj for Go
- throwaway77384 6y agoHow does this compare to something like https://github.com/valyala/fastjson https://github.com/valyala/fastjson ?
- peterohler 6y agoThe fastjson package is fast. Can't disagree but unfortunately it accepts invalid JSON and what you get from the parser is a Value that you can get values from using a Get() function and not a simple go type. To get a simple go type like []interface{} or map[string]interface{} you have to know what the paths are. You can't actually iterate over a map as far as I can tell. Basically the two packages are different in what they provide. OjG returns simple go types and instead of a simple Get() it provides a full JSONPath implementation. Different but both have their uses so depending on a user's needed one might be better than the other.
- svnpenn 6y ago> Both the Ruby Oj and the C parser OjC are the best performers in their respective languages. Um, no, they arent: https://github.com/simdjson/simdjson https://github.com/simdjson/simdjson
- coder543 6y agoAccording to the OjC README: > No official benchmarks are available but in informal tests Oj is 30% faster that [sic] simdjson. Source: https://github.com/ohler55/ojc/blob/master/README.md#benchmarks https://github.com/ohler55/ojc/blob/master/README.md#benchma... If you have benchmarks that show otherwise, that would be great for the discussion here, but your point appears to have already been addressed?
- svnpenn 6y agoThats one person saying it is faster. Without a reproducible test, thats pretty much worthless.
- coder543 6y agoYou don’t see the irony in that statement? You’re one person saying simdjson is faster, and you are also lacking a reproducible test. That doesn’t make your side of the argument particularly compelling either. I’ve never used either library, so I don’t care which one is faster, but I would generally trust the author of the library to have tested the performance at some point more than I would trust someone who isn’t the author to comment on the performance of code they’ve never tested. The author of that library isn’t present to make their case, so you’re the only one of the two who could provide evidence right now. EDIT: the author is present now: https://news.ycombinator.com/item?id=23738512 https://news.ycombinator.com/item?id=23738512
- the-dude 6y agoYour comment would improve without the last sentence.
- coder543 6y agoI was iterating on the comment rapidly after posting to try to make my point in the most neutral way possible, so I’m not sure which “last sentence” you saw, but hopefully it is better now.
- the-dude 6y agoThis is much better, no need for ad hominems. But please keep the editorializing in check too, it is etiquette to mention edits, and continually editting breaks continuity in the discussion. As you pointed out yourself.
- nemothekid 6y ago
- quelsolaar 6y agoI wrote a fast C json parser a few years ago (+600megs/s) And it was an interesting experience. Validating the json roughly halved the performance. The most interesting performance gain was going from using aligned structs to packed byte offsets being accessed using memcpy. It added 20-30%. The overhead of aligning was nothing compared to fewer cache misses. In the end i found that making a truly fast json parser mostly depend on what you parse it to. Like, is the structure read only, and how fast is it to access?
- mister_hn 6y agoWithout validation, all the speed and performance is worthless, if a bad formed JSON can break your application
- beached_whale 6y agoThere are many circumstances where the JSON is always perfect. Having the option to not validate is beneficial
- fnord123 6y agoIf you have so much control over the JSON and performance is a big deal, then there's a big chance you can get rid of the JSON in favor of a more performant format.
- beached_whale 6y agoThat isn't true though. It's the lingua franca of data transfer. Also, why are we giving away money because of this view? We are stuck with JSON for better or worse. I've seen parsers where the same real task, say parse, map, and reduce a 100MB of doubles encoded as JSON. In a decent library it is taking much less than half a second and very little memory, say the size of the result and the memory mapping of the JSON document. In many common libraries it takes multiple seconds while using hundreds/gigs of memory. That means one has to pay for bigger machines with more uptime per task. That is giving money away.
- remorses 6y agoGo is becoming the language to implement json parsers
- mlvljr 6y agoHow dare.
- jeffbee 6y agoThere is a nugget buried at the bottom of this: you will get different output for different kinds of for ... range or for i ... loops. You'd think these are equivalent: for i = range arr { sum += arr[i] } for _, c = range arr { sum += c } ... but the latter is 4 bytes shorter in x86. The compiler takes things literally.
- peterohler 6y agoNice explanation, thanks.
- nautilus12 6y agoI wonder how people that make super specific things like this their whole career make money. I wonder if they are independently wealthy and just do things like this for fun
- peterohler 6y agoMost of us have a day job and do this sort of thing for fun. I know that might seem weird to many but I think you will find most open source developers fall into that category.
- jkeiser 6y agoIt's kind of a "continued learning" requirement, a little. Work doesn't always have the problems (or funding for them) that grow you in the right directions. Open source is also a sort of portfolio for many people, I think. Resumes can only tell you so much; code speaks volumes. It pays, just not directly or immediately. Though my first job in Silicon Valley-sized tech was at Netscape, and I got that specifically after writing some big open source patches for Mozilla. So it can sometimes pay more directly.