7 ms·
Why Go doesn't do this by default: https://github.com/golang/go/issues/10014#issuecomment-91436342 https://github.com/golang/go/issues/10014#issuecomment-91436.
by ferdowsi 5y ago
Why Go doesn't do this by default:
https://github.com/golang/go/issues/10014#issuecomment-91436342 https://github.com/golang/go/issues/10014#issuecomment-91436...
> It seems fine to me for the spec not to guarantee anything about struct field order in memory. The spec doesn't operate at that level.
> That said, no Go compiler should probably ever reorder struct fields. That seems like it is trying to solve a 1970s problem, namely packing structs to use as little space as possible. The 2010s problem is to put related fields near each other to reduce cache misses, and (unlike the 1970s problem) there is no obvious way for the compiler to pick an optimal solution. A compiler that takes that control away from the programmer is going to be that much less useful, and people will find better compilers.
- Ar-Curunir 5y agoWhy not just offer the option to have manual control? That's what Rust does; by default rustc reorders fields, but you can also manually annotate your struct with #[repr(C)], which enforces no reordering.
- roca 5y agoYeah. Automatically optimizing for space gives good cache footprint if you don't know anything about access patterns so it might as well be the default. Also, with generics, there are cases where the optimal field ordering depends on the generic type parameters, so letting the compiler reorder fields on a per-generic-instance basis gives better space utilization than any given source field ordering.
- WatchDog 5y agoSuch a feature would probably require golang to either implement annotations or a new keyword. I'm not sure golang doesn't have annotations, but probably for the same reasons it's taken so long to implement generics, it's a language stuck in the 70's, with a user base and leadership that will fight tooth and nail to try and stay that way.
- Kon-Peki 5y agoGo most definitely allows annotations on structure fields [1]. You say the language is stuck in the 70s, but the GP has a quote from the Go authors saying that having the compiler reorder structure fields is a 1970s problem... Well anyway, mostly it just sounds like the typical Rust enthusiast "How dare you have a strongly-held opinion that differs from my strongly-held opinion". As someone who used to work on a codebase that had to be portable across X86, Alpha, SPARC, and Itanium, paying attention to structure field order quickly becomes a routine matter. It hardly seems worthy of an argument over the merits of a programming language. [1] https://pkg.go.dev/reflect#StructTag https://pkg.go.dev/reflect#StructTag
- adwn 5y ago> Well anyway, mostly it just sounds like the typical Rust enthusiast "How dare you have a strongly-held opinion that differs from my strongly-held opinion" To the contrary, it sounds like "Here's a solution that gives better results in 99% of cases, for reasons X, Y, Z, and better control and guarantees in those other 1% of cases."
- Ar-Curunir 5y agoNo need to be defensive, I was merely suggesting one approach, and didn't even criticize Go anywhere in my comment.
- rsecora 5y agoIts not really a 70s problem, its a solution to a problem if the 1970s. RISC machines raised bus error if not properly aligned to avoid "double" access to the same value. https://en.wikipedia.org/wiki/Bus_error#Unaligned_access https://en.wikipedia.org/wiki/Bus_error#Unaligned_access
- masklinn 5y ago> I'm not sure golang doesn't have annotations It does, the `//go:` stricture is a pragma / annotation. And though I don’t remember any being at the struct-definition level, Go 1.16 added one at the “const” (toplevel var) level so adding it for structs doesn’t seem like an issue.
- jerf 5y agoBecause in practice the Go solution is the other way around: You can optionally decide to worry about it by turning on the linter that detects this. It's not a very different solution in practice. I actually kind of like the idea of a minimal language with a robust linting/static analysis community. You can modularly pick what things you want to worry about. The correct answer for the vast bulk of Go programmers is not to worry about this, and the ones who want to, the tools are readily available and already integrated into a tool that anyone writing serious Go code should already be using.
- kibwen 5y ago> It's not a very different solution in practice. It makes a difference in the presence of generics, since at that point laying out fields efficiently in the face of every possible combination of types is a task that only the compiler can perform. There's nothing that says that it needs to be the default behavior, but if you want efficient space usage then you need more than a lint, you need some way to enable automatic field reordering.
- jerf 5y agoFair enough, I don't think of generics as quite existing yet so I haven't started considering them yet. I suspect in practice we're not going to see a lot of structs with a bajillion generics in them, though. Generics are going to solve the problems the Go developers said they will solve but there's still just enough friction in them (particularly the inability to introduce new types in methods) that I expect it will not be practical to create C++-like libraries of generic things that take generics that take generics as arguments, and in practice, "stick the small number of generic things (most likely one) at the end of the struct" will mostly cover the bases. (I have no problem saying that if you need the n'th degree in performance, you shouldn't have picked Go. I think it has a great bang-for-the-buck ratio, but it definitely does not occupy the "best possible performance" slot.)
- kibwen 5y ago> I suspect in practice we're not going to see a lot of structs with a bajillion generics in them Sure, but note that it only takes a single generic parameter to exhibit this behavior. Consider the original struct definition in the OP: if we imagine that the first field was generic instead of uint8, then the struct has padding only when the type is less than 16 bits in size. No matter where you manually reorder that field, some possible types will still result in padding if the fields are forced to be laid out in order, and it took no more than a single type parameter.
- deleted 5y ago[deleted]
- adonovan 5y agoLook up “false sharing”, the situation where two goroutines that access disjoint sets of fields contend for the same cache line. Compilers are not smart enough to detect it, but if they lay out struct fields in declaration order, the programmer has a way to avoid the problem: by putting the two threads’ fields far apart.
- berkut 5y agoYeah, quite often in HPC you very much do want to be able to control the alignment and packing (even though it's annoying to do), because you'll get caching issues otherwise, depending on the access patterns and size of the data.
- tjoff 5y agoIt seems to follow that you shouldn't use Go in HPC. Which is also fine.
- beltsazar 5y agoHe was basically saying A is an old problem, the new problem is B. However, there's no obvious way to solve B, so we don't solve B. But why not solve A then? Unless.. A is not a problem anymore nowadays. (Is it, though?) But if A is not a problem anymore, he could have just said struct field ordering was an old problem and not a problem anymore in 2010s, without mentioning the other problem. Meanwhile the blog post suggests that struct field ordering is still a problem even in 2020s.
- dllthomas 5y agoThe solution to A will break manual attempts to address B.
- bsuvc 5y agoIf Go tries to automatically solve A, it prevents a developer from solving B because there is no way to control field order. By Go doing nothing, a developer can manually solve both A and B.
- adwn 5y ago> By Go doing nothing, a developer can manually solve both A and B. In Rust, you can annotate a struct declaration with "#[repr(C)]" to prevent the compiler from reordering fields. I don't see why the Go compiler couldn't offer something similar.
- scottlamb 5y agoHow often is "reduce cache misses" that different from "use as little space as possible"? They're basically the same if the struct can be no more than one cacheline wide. When the struct is larger, it's possible you're accessing certain sets of fields together often enough for this to be a useful consideration, but I have no intuition on how common it is. Although it occurs to me that when this happens, switching from arrays of structs to structs of arrays may be a better optimization anyway. fwiw, Rust leaves the ordering unspecified (unless you specify it via #repr[(C)]). Currently it orders to minimize padding. (In theory a future compiler could reorder to minimize cache misses based on a profile or something.) According to https://doc.rust-lang.org/nomicon/repr-rust.html https://doc.rust-lang.org/nomicon/repr-rust.html part of the rationale for reordering was generics. If you have a struct Foo<T, U>, the optimal ordering depends on the size of T and U. The same argument won't apply to Go until 1.18 is released.
- staticassertion 5y ago> How often is "reduce cache misses" that different from "use as little space as possible"? I thought the same exact thing. If I can reduce my struct sizes I can pack N more structs into my cache line. That's almost certainly going to be the best cache-based win.
- deleted 5y ago[deleted]
- tedunangst 5y agoIf your struct is large enough that you care about shaving padding, you probably have hot fields and cold fields and the best cached based win will be arranging them contiguously.
- morelisp 5y agoIf I have a [10000]T I need to shave very little padding before I see some impact, even though both the padded and unpadded T might be relatively small. Specifically in Go, a smaller size can get even lone values into a smaller size class. Saving one byte may save you 384 if it's from 2305 to 2304.
- olliej 5y agoMemory usage is a real problem once you have large amounts of data, of course Go is GC’d and so has a substantial memory hit anyway so I understand weighing that aspect less. The real killer once that’s factored out is cache performance, and that really is a killer: for high performance code you can easily lose double digit %s of perf hit. You can do even worse in, but in the optimal case (a flat array) load predictors and prefetchers get you to only 10-20% hit from cache pressure. This is my recollection from maybe 5 years ago (hell maybe even 10), so it could be worse now.
- deleted 5y ago[deleted]
- jcranmer 5y agoThis is kind of a fake answer to the problem. First off, using less memory is an effective way to reduce cache misses: if you shrink memory by ⅓, that allows you 50% more objects in the same cache size. And this applies to anything--it's the only way to reduce cache misses that is universal. So saying that it's not solving the "real" problem is really a spit-take, because it's a pretty effective way of solving that "real" problem. Suppose you considered cases where a smart ordering could avoid hitting unused cache lines. If a struct is larger than a cache line, it's possible to put co-used values on one cache line and avoid bringing in the other cache lines. But this kind of optimization isn't going to work unless the struct is cache-aligned to begin with--otherwise, your clever ordering is only going to sometimes work and sometimes potentially cause unnecessary multiple cache lines to need to be brought in. As to whether or not cacheline-alignment is a good idea, well, the extra padding will increase memory usage (see point #1), and the potential benefit is going to be limited by how hot or cold field accesses actually are. The other case that comes to mind is false-sharing, which is definitely a real concern. Except, we're talking about reordering struct fields, which means it's false sharing within fields of a struct, and that's a much smaller subset of where false sharing actually occurs--false sharing tends to be more of an issue when you have an array of objects, and you need to make the struct element a multiple of cacheline size to avoid it. The only reasonable cases I can think of off the top of my head are going to involve structs which have intrusive atomic reference counting or some sort of intrusive lock in them--and you can solve both of those cases by making large cacheline-sized versions of those structs that prevent any fields of the outer struct from being stuck on the same cache lines as those data structures. So I rather expect that it is very possible to have a field reordering algorithm that would improve cache misses in all the obvious cases (as I mentioned in point #1) while not preventing the user from having sufficient control to optimize for minimizing cache misses in the rarer cases in the subsequent point.
- rsecora 5y agoIts not only due to cache misses. If the memory access is not aligned a bus fault will be raised in a lot of CPU architectures. Thats the 70s problem described in the documentation