4 ms·
Thanks for posting the update here. I was going to point out that 0.5ns for anything is almost always a sign of benchmarks being optimized away. But even if th
by rsc 2y ago
Thanks for posting the update here. I was going to point out that 0.5ns for anything is almost always a sign of benchmarks being optimized away.
But even if those numbers had been accurate, it is important to think about how they compare to the operations being performed. Taking 4 nanoseconds to check for an error sentinel is not a huge deal when you've spent microseconds or even milliseconds waiting for a network service.
- zachmu 2y agoYou'll never get those 4 nanoseconds of your life back though
- philosopher1234 2y ago4 nanoseconds spent calling high quality functions is 4 nanoseconds well spent
- neonsunset 2y ago4 nanoseconds is how much it takes for unoptimized errors.Is version of this code[0] to execute in C#, instead of 15-17ns in Go. I don't know how long it will take the industry to learn not to use Go for writing databases. [0]: https://gist.github.com/neon-sunset/65a8deeaac745ba159f990c78681f313 https://gist.github.com/neon-sunset/65a8deeaac745ba159f990c7...
- zachmu 2y agoOur database is within spitting distance of MySQL's raw performance and getting faster all the time. Certainly there are some applications where that will not be fast enough. That didn't prevent MySQL from becoming the worlds most widely deployed free database. And yes, postgres is faster and will eventually pass MySQL in deployments, but that's success also has very little to do with its speed.
- vips7L 2y agoIs that AOT or JIT?
- neonsunset 2y agoUpdated the numbers in the gist with more context. This is .NET 9 preview 5 JIT or .NET 8 + DPGO casts and VTable profiling feature (it was tuned, stabilized and is now enabled by default in 9). This would be representative for a long-running DB workload as the one the original blog post refers to, which wants to use JIT rather than AOT. Practically speaking, the example is implemented in a way to make it logically closer to Golang one. However, splitting the interface check and the cast defeats default behavior of .NET 8 JIT's guarded and ILC's exact devirtualization. Ideally, you would not perform type tests and just express it through a single interface abstraction that inherits from IEquatable<T> - this is what regular code looks like and what the compiler is optimized towards (e.g. if ILC can see that only a single type implements an interface, all callsites referring to it are unconditionally devirtualized). Tl;Dr: CPU Freq: 3.22GHz, single cycle: ~0.31ns L1 reference roundtrip lat.: 4 cycles (iirc for Firestorm) Cost of GetValue -> (object?, IError?) + Errors.Is call -- JIT -- .NET 9: 3.24ns value ret, 3.58ns error ret .NET 8: 5.56ns value ret, 5.86ns error ret -- AOT -- .NET 9: 5.72ns value ret, 5.91ns error ret .NET 8: 5.52ns value ret, 5.89ns error ret Note: .NET 9 numbers are not finalized - it's still 5 months away from release.
- vips7L 2y agoVery informative. Thanks. > e.g. if ILC can see that only a single type implements an interface, all callsites referring to it are unconditionally devirtualized). Just like Java then.
- neonsunset 2y agoYeah, it's in the same area as guarded devirtualization done by HotSpot but there are some differences. ILC is "IL AOT Compiler", it targets the same compiler back-end as JIT but each one has its own set of specific optimizations, on top of the ones that do not care about JIT or AOT conditions. JIT relies on "tiered compilation" and DynamicPGO, which is very similar to HotSpot C1/C2, except it does not have interpreter mode - the code is always compiled and data shows that to have better startup latency than e.g. what Spring Boot applications have. AOT on the other hand relies on "frozen world" optimizations[0] which can be both better and worse than JIT's depending on the exact scenario. JIT currently defaults to a single type devirtualization then fallback (configurable). AOT otoh can devirtualize and/or inline up to 3 variants and doesn't have to emit a fallback when not needed because the exact type hierarchy is statically known. There are other differences that stem from architectural choices made long ago, namely, all method calls in C# are direct except those that are explicitly marked as virtual or interface (there are also delegates but eh), naturally, if a compiler sees the exact type in a local scope, no call will be virtual either. In Java, all calls are virtual by default unless stated or proven by compiler otherwise, interface calls there are also more expensive which is why certain areas of OpenJDK have to be this advanced in terms of devirt, and support de- and re-optimization, that .NET's JIT opted not to do - data shows most callsites are predominantly mono and bi-morphic, and the fallback cost is inexpensive in most situations (virtual calls are pretty much like in C++, interface calls use inline caching style approach). Today, JIT, on average, produces faster application code at the cost of memory and startup latency. It also isn't limited by "can only build for -march=x86-64-v2" unlike AOT. In the future, I hope AOT gets the ability to use higher internal compiler limits and spend more time on optimization passes in a way that isn't viable with JIT as it would mean sacrificing its precious throughput. [0]: https://migeel.sk/blog/2023/11/22/top-3-whole-program-optimizations-for-aot-in-net-8/ https://migeel.sk/blog/2023/11/22/top-3-whole-program-optimi...
- rrrix1 2y agoIt may just be me, but I have found greater joy and satisfaction working in Go compared to any other language. I believe this is very important when considering the finite and decreasing time remaining in my life. Sometimes "good enough" performance is perfectly fine when considering a bigger picture.
- neonsunset 2y agoThat's one of the most reasonable ways to look at it, among many voiced here. I agree with it, it's the same reason I enjoy C# (and sometimes F#) - it gives the sense of control, gets out of the way when you want to get things done, and gives powerful tools when you want to push it to the limit. The problem is - Go is not an underdog the way C# is if you look at GitHub statistics, and has reached the escape velocity that gets it picked for all the fun projects, even when they would have been better served by C# which either offers a purpose-built capability or good tools to implement such. And when Go fails that, it will be made work despite its shortcomings, similar to Python and ML. It's very painful to move languages when such move involves the sense of settling for less. I tried a lot and very few felt like an improvement - some would perform better in a particular area but would also have significant shortcomings I'm not comfortable with in areas C# doesn't.
- fl0ki 2y agoIt's a relatively bigger deal for many other potential sources of errors, like parsing a string as a number. Also, even if the caller does not unwrap the error values, producing errors often requires at least one heap allocation. These are all fair tradeoffs for the simplicity and ergonomics of Go's error handling. They rarely matter for performance because the happy critical path of most programs does not produce lots of errors. These are just nuances worth recognizing in a thread specifically about the isolated overheads of errors.
- kiitos 2y agoYep. It should (hopefully!) be self-evident that "hot loops" in which single-digit nanosecond performance differences are meaningful, are (a) exceptional; and, more importantly, (b) not a place where you'd make any kind of function call in the first place!