10 ms·
A new experimental Go API for JSON
- analytically 1y agoBenchmark Analysis: Sonic vs Standard JSON vs JSON v2 in Go https://github.com/centralci/go-benchmarks/tree/b647c45272c7dc371fd4337cb3b6546356d967d1/json https://github.com/centralci/go-benchmarks/tree/b647c45272c7...
- tgv 1y agoIIRC, sonic does JIT, has inline assembly (github says 41%), and it's huge. There's no way you can audit it. If you don't need to squeeze every cpu cycle out of your json parser (and most of us don't; go wouldn't be the first choice for such performance anyway), I'd stick with a simpler implementation.
- jitl 1y agoIt also seems to need 4x the memory
- gethly 1y agoThose numbers look similar to goccy. I used to use it in the past, even Kubernetes uses it as direct dependency, but the amount of issues have been stockpiling for quite some time so I no longer trust it. So it seems both are operating at the edge of Go's capabilities. Personally, I think JSON should be in Go's core and highly optimised simd c code and not in the Go's std library as standard Go code. As JSON is such an important part of the web nowadays, it deserves to be treated with more care.
- tptacek 1y agoWhat does "highly optimized" have to do with whether it's in the standard library? Highly-optimized cryptography is in the standard library.
- ronsor 1y agoNot to mention that Go is never going to put C code in the standard library for anything portable. It's all Go or assembly now.
- dwattttt 1y agoIt's amusing to see assembly considered more portable than C.
- jitl 1y agoNo, portable code is written in Go, not C. Platform specific code is written in ASM.
- deleted 1y ago[deleted]
- pjmlp 1y agoWhich is the right approach, and one of the areas I actually appreciate the work of Go authors. There is nothing special about C, other that its historical availability after UNIX's free beer came to be. Any combination of high level + Assembly is enough.
- dilyevsky 1y agoPreviously Go team has been vocal about sacrificing performance to keep stdlib idiomatic and readable. Guess the crypto packages are the exception because they are used heavily by Google internally and json and some others (like say image/jpeg which had crap performance last time i checked) are not. Edit: See: https://go.dev/wiki/AssemblyPolicy https://go.dev/wiki/AssemblyPolicy
- karel-3d 1y agosonic uses a clang-generated ASM, built from C (transformed from normal clang-generated ASM to "weird" go ASM via python script)... I don't think this will be in standard library.
- nasretdinov 1y agoGo doesn't yet have native SIMD support, but it actually might in the future: https://github.com/golang/go/issues/73787 https://github.com/golang/go/issues/73787 I think when it's introduced it might be worth discussing that again. Otherwise providing assembly for JSON of all packages seems like a huge maintenance burden for very little benefit for end users (since faster alternatives are readily available)
- kristianp 1y agoInteresting that the proposal is for low-level intrinsics as well as a non-processor-specific api: > Our plan is to take a two-level approach: Low-level architecture-specific API and intrinsics, and a high-level portable vector API. The low-level intrinsics will closely resemble the machine instructions (most intrinsics will compile to a single instruction), and will serve as building blocks for the high-level API.
- CamouflagedKiwi 1y agoAgreed. goccy has better performance most times but absolutely appalling worst-case performance which renders it unacceptable for many use cases - in my case even with trusted input it took effectively eternity to decode it. It's literally a quadratic worst case, what's the point of having a bunch of super clever optimisations if the big-O performance is that bad. Sonic may be different but I'm feeling once bitten twice shy on "faster" JSON parsers at this point. A highly optimised SIMD version might be nice but the stdlib json package needs to work for everything out there, not just the cases the author decided to test on, and I'd be a lot more nervous about something like that being sufficiently well tested given the extra complexity.
- godisdad 1y ago> As JSON is such an important part of the web nowadays, it deserves to be treated with more care. There is a case to be made here but Corba, SOAP and XML-RPC likely looked similarly sticky and eternal in the past
- mdaniel 1y agoI hear you, but I am not aware of anyone that tried XMLHttpRequest.send('<s:Envelope xmlns:s...') or its '<methodCall>' friend from the browser. I think that's why they cited "of the web" and not "of RPC frameworks"
- pjmlp 1y agoNo, because that was server's job on the endpoint.
- int_19h 1y agoI don't recall either CORBA or SOAP ever seeing enough penetration to look "eternal" as mainstream tech goes (obviously, and especially with SOAP, there's still plenty of enterprise use). Unlike XML and JSON.
- pjmlp 1y agoThey surely were, for anyone doing enterprise during the 2000's. We had no plans to change to something else.
- int_19h 1y agoI recall a lot of talk about CORBA in early 00s, but I don't think I've actually ever seen it used anywhere outside of Gnome. By late 00s, even the talk was more along the lines of it being legacy tech.
- pjmlp 1y agoSeveral Nokia Networks products were based on CORBA, running on HP-UX, in a mix of C++ and Perl. Eventually migrated to Java EE, also taking advantage of CORBA compatibility.
- ForHackernews 1y agoThe fact that JSON is used so commonly for web stuff seems like an argument against wasting your time optimizing it. Network round trip is almost always going to dominate. If you're pushing data around on disk where the serialization library is your bottleneck, pick a better format.
- catlifeonmars 1y agoYou’re assuming request-response round trip between each call to encode/decode. Streaming large objects/NDJSON would still have serialization bottleneck. (See elasticsearch/opensearch for a real life use case) But in that case your last point still stands: pick a better format
- gethly 1y agoThere is no better human-readable format. I looked. The only alternative i considered was Amazon Ion but it proved to bring no additional value compared to json.
- kbolino 1y agoThis is an interesting perversion of Amdahl's law. Yes, if you are looking at a single request-response interaction over the Internet in isolation and observing against wall clock time, the time spent on JSON (de-)serialization (unless egregiously atrocious) will usually be insignificant. But that's just one perspective. If we look at CPU time instead of wall clock time, the JSON may dominate over the network calls. Moreover, in a language like Go, which can easily handle tens to hundreds of thousands of parked green threads waiting for network activity, the time spent on JSON can actually be a significant factor in request throughput. Even "just" doubling RPS from 10k to 20k would mean using half as much energy (or half as much cloud compute spend etc.) per request. Changing formats (esp to a low-overhead binary one) might yield better performance still, but it will also have costs, both in time spent making the change (which could take months) and adapting to it (new tools, new log formats, new training, etc.).
- ForHackernews 1y agoIf you're optimizing for energy wasted serving your website you could stop sending 10 megs of garbage javascript on page load.
- deleted 1y ago[deleted]
- kiitos 1y agofirst of all, that doesn't exercise JSON v2 at all, afaict second of all, sonic apparently uses unsafe to (unsafe-ly) cast byte slices to strings, which of course is gonna be faster than doing things correctly, but is also of course incomparable to doing things correctly like almost all benchmark data posted to hn -- unsound, ignore
- kristianp 1y agoJust using the GOEXPERIMENT=jsonv2 compiler flag changes the underlying implementation if you don't change any code. You're still using the less correct and efficient API though.
- Thaxll 1y agoAnd Sonic with its "cutting edge" optimization is still slower than std Json on arm64 with basic use cases. It shows that JIT, simd, low level code comes at cost of maintenance for all platform. https://github.com/bytedance/sonic/issues/785 https://github.com/bytedance/sonic/issues/785
- gethly 1y agoThis V2 is still pushing forward the retarded behavior from v1 when it comes to handling nil for maps, slices and pointers. I am so sick and tired of this crap. I had to fork the v1 to make it behave properly and they still manage to fuck up completely new version just as well(by pushing omitempty and ignoring omitnil behavior as a standalone case) which means I will be stuck with the snale-pace slow v1 for ever.
- ycombinatrix 1y agoWhat is your preferred behavior for a nil map/slice? Feels weird that it doesn't map to null.
- gethly 1y agoWhen you are unmarshaling json, empty map/slice is something completely different than a null or no value present, as you are losing intent of the sender, in case of JSON REST. For example, if my intent is to keep the present value, I will send {"foo": 1} or {"foo": 1, "bar": null} as null and no value has the same meaning. On the other hand, I might want to change the existing value to empty one and send {"foo": 1, "bar": []}. The server must understand case when I am not mutating the field and when I am mutating the field and setting it to be empty. On the other side, I never want to be sending json out with null values as that is waste of traffic and provides no information to the client, ie {"foo": 1, "bar": null} is the same as {"foo": 1}. Protocol buffers have the exact same problem but they tackle it in even dumber way by requiring you to list fields from the request, in the request's special field, which you are mutating as they are unable to distinguish null and no value present and will default to empty value otherwise, like {} or [], which is not the intent of the sender and causes all sort of data corruption. PS: obviosly this applies to pointers as a whole, so if i have type Request struct { Number *int `json:"number"} then sending {} and {"number": null} must behave the same and result in Result{Number nil}
- deleted 1y ago[deleted]
- 1y ago
- rjrodger 1y agonull != nil !!! It is good to see some partial solutions to this issue. It plagues most languages and introduces a nice little ambiguity that is just trouble waiting to happen. Ironically, JavaScript with its hilarious `null` and `undefined` does not have this problem. Most JSON parsers and emitters in most languages should use a special value for "JSON null".
- pjmlp 1y agoFixed in 1976 by ML, followed up by Eiffel in 2005, but unfortunately yet to be made common.
- afiori 1y agoNull and undefined are fine imho with a sort of empty/missing semantics (especially since you mostly just care to == them) I have bigger issues to how similar yet different it is to have an undefined key and a not-defined key, I would almost prefer if obj['key']=undefined was the same as delete obj['key']
- coldtea 1y ago>Over time, packages evolve with the needs of their users, and encoding/json is no exception No, it's an exception. It was badly designed from the start - it's not just that people's json needs (which hardly changed) outgrew it.
- bayindirh 1y agoA bad design doesn't invalidate the sentence you have quoted. Over time, it became evident that the JSON package didn't meet the needs of its users, and the package has evolved as a result. The size of the evolution doesn't matter.
- pcthrowaway 1y agoIt's true that packages (generally) evolve with the needs of their users. It's also true that a json IO built-in lib typically wouldn't be so poorly designed in the first release of a language, that it would immediately be in need of maintenance.
- bayindirh 1y ago> immediately JSON library released with Go 1, in 2012. This makes the library 13 years old [0]. If that's immediate, I'm fine with that kind of immediate. [0]: https://pkg.go.dev/encoding/json@go1 https://pkg.go.dev/encoding/json@go1
- pcthrowaway 1y agoIn need of maintenance and having received maintenance are two different things
- coldtea 1y ago"immediately be in need of maintenance" means it needed this update 13 years ago.
- 1y ago
- sroerick 1y agoCould somebody give a high level overview of this for me, as not a godev? It looks like Go JSON lib has support to encode native go structures in JSON, which is cool, but maybe it was bad, which is not as cool. Do I have that right?
- eknkc 1y agoGo already has a JSON parser and serializer. It kind of resembles the JS api where you push some objects into JSON.stringify and it serializes them. Or you push some string and get an object (or string etc) from JSON.parse. The types themselves have a way to customize their own JSON conversion code. You could have a struct serialize itself to a string, an array, do weird gymnastics, whatever. The JSON module calls these custom implementations when available. The current way of doing it is shit though. If you want to customize serialization, you need to return a json string basically. Then the serializer has to check if you actually managed to return something sane. You also have no idea if there were some JSON options. Maybe there is an indentation setting or whatever. No, you return a byte array. Deserialization is also shit because a) again, no options. b) the parser has to send you a byte array to parse. Hey, I have this JSON string, parse it. If that JSON string is 100MB long, too bad, it has to be read completely and allocated again for you to work on because you can only accept a byte array to parse. New API fixes these. They provide a Decoder or Encoder to you. These carry any options from top. And they also can stream data. So you can serialize your 10GB array value by value while the underlying writer writes it into disk for example. Instead of allocating all on memory first, as the older API forces you to. There are other improvements too but the post mainly focuses on these so thats what I got from it (I havent tried the new api btw, this is all from the post so maybe I’m wrong on some points)
- stackedinserter 1y agogjson/sjson is probably for you if you need to work with 100MB JSONs.
- trimethylpurine 1y agoThis is cool. I wouldn't have thought to use Go for stuff that size.
- tibbe 1y ago> Since encoding/json marshals a nil slice or map as a JSON null How did that make it into the v1 design?
- rowanseymour 1y agoI had a back and forth with someone who really didn't want to change that behavior and their reasoning was that since you can create and provide an empty map or slice.. having the marshaler do that for you, and then also needing a way to disable that behavior, was unnecessary complexity.
- binary132 1y agohow is a nil map not null? It certainly isn’t a zero-valued map, that would be {}.
- atombender 1y agoThe zero value of a map is indeed nil in Go: This prints true (https://go.dev/play/p/8dXgo8y2KTh https://go.dev/play/p/8dXgo8y2KTh): var m map[string]int println(m == nil)
- binary132 1y agoOk, true!
- materielle 1y agoIt should be marshaled into {} by default, with a opt-out for special use cases. There’s a simple reason: most JavaScript parsers reject null. At least in the slice case.
- tubthumper8 1y agoNot sure what you mean here by "most JavaScript parser rejects null" - did you mean "JSON parsers"? And why would they reject null, which is a valid JSON value? It's more that when building an API that adheres to a specification, whether formal or informal, if the field is supposed to be a JSON array then it should be a JSON array. Not _sometimes_ a JSON array and _sometimes_ null, but always an array. That way clients consuming the JSON output can write code consuming that array without needing to be overly defensive
- h1fra 1y agoI still don't get how a common thing like JSON is not solved in go. How convoluted it is to just get a payload from an api call compared to all languages is baffling
- Thaxll 1y agoYou should read that, it's still relevant: https://seriot.ch/projects/parsing_json.html https://seriot.ch/projects/parsing_json.html
- deleted 1y ago[deleted]
- 9rx 1y ago> I still don't get how a common thing like JSON is not solved in go. Given that it is not even yet solved in its namesake language, Javascript, that's not saying much.
- breakingcups 1y agoIf/once this goes through, I wonder what the adoption is going to be like now that all LLMs still only have the v1 api in their corpus.
- h4ch1 1y agoHopefully people will remember documentation exists once errors start popping up and refer to it.
- afdbcreid 1y agoThis is the second time a v2 is released to a package in the Go's standard library. Other ecosystems are not free of this problem. And then people complain that Rust doesn't have a batteries-included stdlib. It is done to avoid cases like this.
- oncallthrow 1y agoWow, two whole times in 19 years? That sounds terrible. Yes, we should definitely go with the Rust approach instead. Anyway, I'd better get back to figuring out which crate am I meant to be using...
- ncruces 1y agoThat has its own downsides, though. Both v1 packages continue work; both are maintained. They get security updates, and were both improved by implementing them on top of v2 to the extent possible without breaking their respective APIs. More importantly: the Go authors remain responsible for both the v1 and v2 packages. What most people want to avoid with a "batteries included standard library" (and few additional dependencies) is the debacle we had just today with NPM. Well maintained packages, from a handful of reputable sources, with predictable release schedules, a responsive security team and well specified security process. You can't get that with 100s of independently developed dependencies.
- kiitos 1y agoI'm not sure how this is a problem, and I'm very sure that even in the presence of this "problem" it is far better for a language to have a batteries-included stdlib than to not
- skywhopper 1y agoTwo v2s in 15 years seems pretty good given the breadth of the stdlib.
- jitl 1y agoI’d rather have 2 jsons in the stdlib after 15 years than 0 jsons in the stdlib
- 1y ago
- phoenixhaber 1y agoI will say this and I feel it's true. Dealing with JSON in Go is a pain. You should be able to write json and not have to care about the marshalling and the unmarshalling. It's the way that serde rust behaves and more or less every other language I've had to deal with and it makes managing this behavior when there's multiple writers complicated.
- dmoy 1y ago> serde rust That does look a lot cleaner. I was just grumbling about this in golang yesterday (yaml, not json, but effectively the same problem).
- lsaferite 1y agoI work in go every day and generally enjoy it. The lack of tagged unions of some sort in go makes things like polymorphic json difficult to handle. It's possible, but requires a ton of overhead. Rust with enums mixed with serde makes this trivial. Insanely trivial.
- romantomjak 1y agoAre you referring to the json macro that allows variable interpolation? Doing that will void type safety. Might be useful in dynamic languages like Python but I wouldn’t want to trade type safety for some syntactic sugar in Go
- curtisszmania 1y ago[dead]
- physicles 1y agoLove seeing meaningful stdlib improvements. I just ran our full suite of a few thousand unit tests with GOEXPERIMENT=jsonv2 and they all passed. (well, one test failed because an error message was changed, but that's on us) I'm especially a fan of breaking out the syntactic part into into its own jsontext package. It makes a ton of sense, and I could see us implementing a couple parsers on top of that to get better performance where it really matters. I wish they would take this chance to ditch omitempty in favor of just the newly-added omitzero (which is customizable with IsZero()), to which we'll be switching all our code over Real Soon Now. The two tags are so similar that it takes effort to decide between them.
- kbolino 1y agoI think the "omitempty" tag name might be too tarnished to keep around, but I think the distinction made between its redefined meaning and the new "omitzero" tag in v2 is quite useful: - "omitempty" will omit an object field after encoding it, according to its JSON value - "omitzero" will omit an object field before encoding it, according to its Go value The former is particularly useful when you are dealing with foreign types that don't implement IsZero (yet) or implement it in an inappropriate way for how you're using it. You could, of course, write a wrapper type, but even when you can use struct embedding to make the wrapper less painful, you still have to duplicate all of the constructors/factories for that type, and you have to write the tedious code to do the conversions somewhere.
- donatj 1y agoI'm coming in a little hot and contrarian. I've been working with the Go JSON library for well over a decade at this point, since before Go 1.0, and I think v1 is basically fine. I have two complaints. Its decoder is a little slow, PHP's decoder blows it out of the water. I also wish there was an easy "catch all" map you could add to a struct for items you didn't define but were passed. None of the other things it "solves" have ever been a problem for me - and the "solution" here is a drastically more complicated API. I frankly feel like doing a v2 is silly. Most of the things people want could be resolved with struct tags varying the behavior of the existing system while maintaining backwards compatibility. My thoughts are basically as follows The struct/slice merge issue? I don't think you should be decoding into a dirty struct or slice to begin with. Just declare it unsupported, undefined behavior and move on. Durations as strings? Why? That's just gross. Case sensitivity by default? Meh. Just add a case sensitivity struct tag. Easy to fix in v1 Partial decoding? This seems so niche it should just be a third party libraries job. Basically everything could've been done in a backwards compatible way. I feel like Rob Pike would not be a fan of this at all, and it feels very un-Go. It goes against Go's whole worse is better angle.
- dematz 1y agoI like the "does the problem justify the solution's complexity" question. The deserialization performance improvement seems like an actually important benefit though. Also https://antonz.org/go-json-v2/#marshalwrite-and-unmarshalread https://antonz.org/go-json-v2/#marshalwrite-and-unmarshalrea... not completely sure but maybe combining dec := json.NewDecoder(in) dec.Decode(&bob) to just json.UnmarshalRead(in, &bob) is nicer...mostly the performance benefit though
- dematz 1y agoAlthough actually, for streaming maybe it would still be 2 lines but from jsontext.. dec := jsontext.NewDecoder(in) json.UnmarshalDecode(in, &bob)
- saghm 1y ago> Durations as strings? Why? That's just gross > It goes against Go's whole worse is better angle One could almost say that durations as strings is...worse.
- drej 1y agoPlease do run this on your own workloads! It's fairly easy to set up and run. I tried it a few weeks ago against a large test suite and saw huge perf benefits, but also found a memory allocation regression. In order for this v2 to be a polished release in 1.26, it needs a bit more testing.