5 ms·
> This is what generics are for in most languages. That goes exactly opposite of this optimization, and is just entirely orthogonal. > It's also super disappo
by bfgoodrich 4y ago
> This is what generics are for in most languages.
That goes exactly opposite of this optimization, and is just entirely orthogonal.
> It's also super disappointing that in 2022 we're still manually marking 10- line functions used in one place and called from one site as inline.
You misunderstand the change.
In a nutshell the prior version was a polymorphic function pointer that varied by the sort type.
(*sort_ptr)(data);
No compiler can inline this, no matter how aggressively you turn up the optimizations.
Their change adds specialized cases.
if (dataType == PG_INT32) {
sort_int32(data);
} else if (dataType == PG_INT64) {
sort_int64(data);
} ... {
} else {
(*sort_ptr)(data);
}
This enabled the inlining, and while it's no harm to add the inline specifier if you want to be very overt in your intentions, most compilers would inline it regardless in most scenarios.
- vlovich123 4y agoIn c++, the sort algorithm is polymorphic and inlined because it automatically generates those sort variants for you. Op is not trying to say “the code as is” but “written in c++ style”. I don’t think they misunderstood anything (and their beef with seemingly unprofiled use of always inline attribute is probably correct too)
- bfgoodrich 4y ago> I don’t think they misunderstood anything The code could have made the type-specific sort macros. They could have literally written the code blocks in the actual sort function. Instead they decided to make it separate functions and mark it inline because that is their overt intention and they thought it was cleaner. What is the problem? Honestly, aside from noisy, useless griping, what is possibly the problem with that? Their implication seems to be that inline shouldn't be necessary. And the truth is that it almost certainly isn't necessary, and any optimizing compiler would inline regardless. But clarifying their intentions is precisely what developers should do: "We're moving it here for organizational/cleanliness purpose, but it is intentionally written this way to be inlined because that is the entire point of this optimization". As to the "unprofiled" use, their specific optimization was because they know the time being wasted in function calls, and the optimization was to avoid that.
- macdice 4y ago(Author of the code here) Yeah :-)
- samhw 4y ago>> This is what generics are for in most languages. > That goes exactly opposite of this optimization, and is just entirely orthogonal. What does that mean? Also, how is it "exactly opposite" and "orthogonal"? Maybe let's get rid of the mixed metaphors and speak in concrete terms. In your comment below ("why would you want the compiler to do that? you can write out all the monomorphizations yourself!") I'm not totally certain that you understand the point of generics. You can always write out each case yourself. Incidentally that also applies to your precious inlining. But it's a tool to make programmers more productive.
- deleted 4y ago[deleted]
- wahern 4y ago> In a nutshell the prior version was a polymorphic function pointer that varied by the sort type. > (*sort_ptr)(data); > No compiler can inline this, no matter how aggressively you turn up the optimizations. LTO can inline these just fine, provided the object files contain the necessary info (i.e. the IR or assembly from static archives, which is how -flto works for clang and GCC). I even tested this using a copy[1] of OpenBSD's qsort implementation a few weeks ago, verifying that both the sort algorithm itself as well as the comparator were inlined, just as would happen with C++ template'd vector sorting. [1] I had to use a copy of qsort because the compiled libc qsort lacks the accompanying info necessary for LTO. But in theory libc and shared libraries in general could ship critical routines with the necessary metadata required for LTO optimizations. (IOW, a cross between a dynamic and static library.) Fundamentally this is just a toolchain issue, not a language issue. AFAIU, Rust code works much the same way, and its comparator functions tend to be inlined because Rust effectively compiles using LTO (and heavily relies on these implicit optimizations), except that because all Rust code is compiled statically, the necessary IR is always available. The same would be true of Go, but the Go compiler doesn't optimize as heavily. FWIW, AFAIU the C++ standard doesn't guarantee a templated sort comparator would get inlined, either, even if you're using a lambda. But this is optimization is extremely likely when the sort algorithm itself is template-expanded and the comparator is simple--as almost all would be.