7 ms·
Gobs of data (2011)
- buro9 3y agothis is quite old, so I'm curious about what triggered it being posted again, has something happened / changed?
- jstanley 3y agoJust because you already knew it all doesn't mean everyone else did. I hadn't seen it before. Sometimes even when something was posted a few years ago some people just haven't seen it yet.
- sudhirj 3y agoTen thousand people, to be exact https://xkcd.com/1053/ https://xkcd.com/1053/
- blowski 3y agoIt's an entirely reasonable question to ask "is there any specific as to why this is being posted today?". If the answer is no, that's fine, but there may be extra context that is interesting and not obvious.
- ash 3y agoI've posted it because I'm always on the lookout for simple solutions for complex problems, and especially for how these solutions are designed. The post describes the design process well. Also Rob Pike is a great technical writer. Another example of his style is "Effective Go": https://go.dev/doc/effective_go https://go.dev/doc/effective_go
- buro9 3y agoyup, and if people are looking for usage I just found a gist that shows how gob handling can be useful (writing to cache that allows the reading back to be castable into the correct structs) https://gist.github.com/pioz/ca5b7a11200f54afbd76dee7acbcc067 https://gist.github.com/pioz/ca5b7a11200f54afbd76dee7acbcc06...
- azaras 3y agoI did not know it, but I think so few changes from the proto-buffer that it is a waste of time.
- bheadmaster 3y agoNote that this was written in 2011, while the first mention of "proto3" in protobuf repository was in 2014. So this blogpost probably influenced the development of proto3, which fixed many issues of proto2 (which is referred to as just "protocol buffers" in the blogpost).
- icholy 3y agoEh, I like Go and respect Rob Pike, but I seriously doubt gob had any impact on the proto3 design
- losvedir 3y agoInteresting. I wonder to what extent it's found use at Google over this past decade. There are advantage to being language-specific, but a lot of disadvantages, as well (speaking as someone who recently had to write some Elixir code to unmarshal a Ruby object...). It seems hard to introduce this since you're forcing all communicating services to be Go-based, which is kind of contrary to the independence that microservices usually affords you. Some of the benefits are simply design goals (e.g., top level arrays) which could also be done in a language-independent protocol. And even performance questions probably could. Like, Cap'n Proto I think is designed so that users of the protocol don't have to serialize/deserialize the data, right? They just pass it around and work with it directly. I can see Rob Pike being frustrated with Protocol Buffers at Google, and I don't begrudge anyone for taking a big shot like this, but I wonder if he's found any success with it.
- lifthrasiir 3y agoYeah, after years of dealing with language-specific serialization formats---and inadvertently learning internals of them (including Go gob, Python pickle and PHP serialize), I'm over. And gob is not even a schematic serialization format (i.e. not only you don't need to define a schema beforehand, you can't). There is some interesting idea, but that's all. Use a well-known schemaless serialization format with some extensibility [1] if you really need. [1] Maybe there was no suitable one when Go was first created. Nowadays I believe CBOR is the best format for this job.
- packetlost 3y agoI'm in the same boat. Not to mention security concerns that often crop up in (interpreted) language specific deserialization (I'm looking at you, pickle, thinly veiled `eval()`). I agree that CBOR should generally be the serialization tool of choice for self-describing data (ie. in places where you might otherwise choose JSON). And if your language of choice doesn't have a CBOR lib, CBOR is fairly easy to implement and writing a encoder/decoder is very fun! I recently completed my implementation for the Gerbil Scheme language last week [0]. [0]: https://github.com/chiefnoah/gerbil-cbor https://github.com/chiefnoah/gerbil-cbor
- 3y ago
- dmi 3y ago> If all you want to send is an array of integers, why should you have to put it into a struct first? If you're sure that's all you'll ever have to do, then sure. But unless you're 100% certain that the protocol will never evolve further, having a more complex structure allows it to change in a gradual way.
- lsaferite 3y agoIt was clear, from the post, that they were saying, "If all I need is a simple array, why should I be required to wrap it in a struct?" The whole point (from the post) being that protobuf required structs but gob allowed simpler types _in addition_ to structs.
- Thorrez 3y agodmi knows that. dmi was saying that even if the encoding scheme allows encoding simpler types, it's often not smart to use that functionality, because you won't be able to evolve the format in the future. If you encode a message instead of a simple type, you'll be able to evolve it later as you add more features to your program. Note that even protobufs, which doesn't allow encoding simple types at the top level, still has this debate when deciding whether to encode an array of simple types (inside a struct) or an array of structs (inside a struct). And Google's guidance is to use an array of structs if more data might be needed in the future: >However, if additional data is likely to be needed in the future, repeated fields should use a message instead of a scalar proactively, to avoid parallel repeated fields. https://google.aip.dev/144 https://google.aip.dev/144 >// Good: A separate message that can grow to include more fields https://protobuf.dev/programming-guides/api/#order-independence-repeated-fields https://protobuf.dev/programming-guides/api/#order-independe...
- orf 3y agoall I need _right now_ is a simple array Nobody knows the future, and preparing for the future is a huge part of software engineering. Sending top-level arrays instead of sending them inside a struct is never the right way.
- assbuttbuttass 3y agoGob is a great serialization format! It's super easy to use, and supports go native types (kind of like Python's pickle). For a recent project, I needed a simple key-value store. I was evaluating using a full RDBMS, but I ended up just putting gob files in a directory.
- jeffrallen 3y agoI used gob for my first client/server Go program, which was a "make one of something you know about to throw away" new language experiment. It worked, but I quickly turned away from it, because it would never be cross platform. I saw gob more as an experiment that the Go team used to check the reflect package's usability. (Which sucks anyway, by the way.) I'm surprised it's still in the stdlib. I would have guessed it would have been removed for Go 1.0, because it was already clear then that it was not suitable for anything more experiments.
- sebstefan 3y ago>[Required fields are] also a maintenance problem. Over time, one may want to modify the data definition to remove a required field, but that may cause existing clients of the data to crash. Okay, but would you rather have it crash or allow for a program to run on the wrong data? Especially if you do that and then say that everything has zero as a default value. The question remains whether the serialization format should be taking care of that, or a round of parsing later on with a schema on the side; but if you do the former without the latter you're setting yourself up for deployment nightmares
- tomohawk 3y agoIf you want to deal with the crash and justify why the system went down because you were more correct than the other guys, then sure. Protocols often represent an interface between organizations. Especially when that is the case, you want to be as charitable as possible when accepting input, because getting any issues resolved may very well require getting the two organizations to agree. Also, as things change over time, an overly strict interpretation when receiving packets will require unnecessary rework in the future, and possibly down time or lost business. When dealing with protocols, it's generally best to be strict when emitting packets and as tolerant as possible when accepting them.
- sebstefan 3y agoThat's the motto for browsers and I agree with it in context, but if it's something you control (like services of a distributed application) then not really. You can just make sure the versions match during deployment and save yourself some debugging headaches Not if it's something sensitive either, where maybe crashing is preferable to running the wrong way
- mst 3y agoLargely agree, with the addendum that it's a really good idea to collect metrics as to how much tolerance your code has been required to show. Whether you need to present those metrics to the sender and ask them to tweak their emissions or simply keen an eye on them is situation dependent, but having them at all is definitely in the "future you will thank current you later" ... and I will absolutely confess that current me has cursed past me for not doing so on more than one occasion, and I can only hope I remember more often in the future ;)
- art_vandalay 3y agoI forgot Go was still around. Thanks for reminding me.
- broken_broken_ 3y agoJust finished removing this encoding in our production services. It panics on malformed input which is a no go for us since high availability is really important for us, and it showed quite a lot in the performance and memory profiles (roughly 5 times the time and memory as doing the same with JSON). The code was converting some data to gob, and storing it in the database for later. We now just do the same but in json, it’s human readable and Postgres validates that the data is valid JSON. And unmarshaling it does not panic.
- jjtheblunt 3y agoHave you tried the superset of json from AWS? https://amazon-ion.github.io/ion-docs/ https://amazon-ion.github.io/ion-docs/
- abtinf 3y agoI've been considering adopting the gob package. I haven't used it before, so I only know what's in the docs -- and all of your claims are surprising to me. Could you share more information? How is it possible that they were getting malformed input? This was happening in go-to-go communication, or was there some kind of cross-language interop? Any idea why the performance was so much slower than JSON in your case? The technique described in the OP would seem to make that impossible. Do you think it's possible the database column type or collation was somehow affecting the gob?
- broken_broken_ 3y agoThe column type was bytea (basically blob) so it should be stored as is by the database. The profiling showed the hotspots in the gob package directly. The docs explicitly mention that invalid input will make it panic and that can be confirmed by reading the code or fuzzing the input. From my understanding, there is no compile time schema so everything is done with runtime reflection and that is bound to not be super fast. Granted, JSON is the same on paper, I would guess that the JSON package had more eyes on it and optimizations. In our case, everything was using JSON except this one component due to some historical oddity so it was also a win in terms of simplifying.
- 3y ago
- jerf 3y agoFWIW, this isn't used much by the community. Being a standard library package it still get some use of course, but for comparison, encoding/gob shows about 22.5K imports [1] to encoding/json's nearly 800K, and whereas you can see in the JSON search an ecosystem of JSON libraries, gob is basically just gob. Calling it "dead" just invites a tedious thread about what the definition of "dead" is, so I won't, I'll just sort of imply it in this sentence without actually coming out and saying it in a clear manner. I would generally both A: recommend against this, not necessarily as a dire warning, just, you know, a recommendation and B: for anyone who is perturbed by the idea of this existing, just be aware that it's not like this package has embedded itself into the Go ecosystem or anything. [1]: https://pkg.go.dev/search?q=gob https://pkg.go.dev/search?q=gob [2]: https://pkg.go.dev/search?q=json https://pkg.go.dev/search?q=json
- emmanueloga_ 3y agoSomeone made a benchmark of serialization libraries in go [1], and I was surprised to see gobs is one of the slowest ones, specially for decoding. I suspect part of the reason is that the API doesn't not allow reusing decoders [2]. From my explorations it seems like both JSON [3], message-pack [4] and CBOR [5] are better alternatives. By the way, in Go there are a like a million JSON encoders because a lot of things in the std library are not really coded for maximum performance but more for easy of usage, it seems. Perhaps this is the right balance for certain things (ex: the http library, see [6]). There are also a bunch of libraries that allow you to modify a JSON file "in place", without having to fully deserialize into structs (ex: GJSON/SJSON [7] [8]). This sounds very convenient and more efficient that fully de/serializing if we just need to change the data a little. -- 1: https://github.com/alecthomas/go_serialization_benchmarks https://github.com/alecthomas/go_serialization_benchmarks 2: https://github.com/golang/go/issues/29766#issuecomment-454926474 https://github.com/golang/go/issues/29766#issuecomment-45492... -- 3: https://github.com/goccy/go-json https://github.com/goccy/go-json 4: https://github.com/vmihailenco/msgpack https://github.com/vmihailenco/msgpack 5: https://github.com/fxamacker/cbor https://github.com/fxamacker/cbor -- 6: https://github.com/valyala/fasthttp#faq https://github.com/valyala/fasthttp#faq -- 7: https://github.com/tidwall/gjson https://github.com/tidwall/gjson 8: https://github.com/tidwall/sjson https://github.com/tidwall/sjson
- dang 3y agoNot gobs of comments but discussed at the time: Gobs of data - https://news.ycombinator.com/item?id=2365430 https://news.ycombinator.com/item?id=2365430 - March 2011 (2 comments)
- zgiber 3y agoIt may not be a good tool for communicating between services implemented in different languages. But i’d happily use it to save stuff to disk where database is overkill.