4 ms·
I was all set to defend Scheme, which has several extremely high performance implementations—and then realized you probably just meant Smalltalk, an entirely di
by brians 2y ago
I was all set to defend Scheme, which has several extremely high performance implementations—and then realized you probably just meant Smalltalk, an entirely different language, and picked the wrong 1970s minimalist language starting with S.
- gizmo 2y agoOops, fixed that. (But as far as I know even performance-oriented Scheme compilers like Bigloo don't generate very efficient code nor do they auto-vectorize. I don't think there is any Scheme compiler that is in the same league as clang performance wise, but I'm happy to be corrected on that.)
- chuckadams 2y agoThe unfortunately-named Stalin compiler boasted pretty good performance of the output (the compiler itself not so much). Hasn't been maintained since R4RS days tho, so probably not useful for any real-world code.
- gizmo 2y ago15 years ago CPUs cared much more about branches and much less about cache locality and threading. Nowadays it's actually really hard to shovel data into the CPU as quickly as it can churn through it. The kind of optimizations talked about here[1] won't get you to even 1% of the performance your CPU is capable of. A Ryzen 5950X can do something like 5ghz * 4 IPC * 16 cores = 300 billion ops per second. Add SIMD on top of that. It's insane. [1] https://cstheory.stackexchange.com/questions/9765/the-stalin-compiler-brutally-optimizes-but-how https://cstheory.stackexchange.com/questions/9765/the-stalin...
- sfn42 2y agoIs this something I should know about as a C# programmer or is it largely handled by the language? I usually don't worry about low level optimization, focusing mostly on writing code that is reasonably efficient with regards to time and space complexity. I know there's a lot of gain to be had from minimizing allocations and such, but for most things it seems pointless to worry about.
- chuckadams 2y agoIt's something you should rely on the language to optimize, but it always helps to not allocate frivolously if you can help it. I'm talking about using StringBuilder rather than concatenation, avoiding unnecessary boxed types, etc. Pooling every last thing doesn't do performance any favors mind you -- the tenured generation is expensive to collect, whereas eden is just a pointer bump.
- neonsunset 2y agoGenerally speaking, transient buffers today, in performance sensitive code in NET, follow the pattern of T[]? toReturn = null; var buffer = length <= threshold ? stackalloc T[threshold] : (toReturn = ArrayPool<T>.Shared.Rent(length)); /* logic */ if (toReturn != null) ArrayPool<T>.Shared.Return(toReturn); either directly or via a buffer-like type that does it behind the scenes. ArrayPool<T>.Shared is generally well-behaved in terms of GC, and small lengths will not even hit it, being practically free. The amortized cost of this is substantially lower than allocating such arrays and then throwing them away: stackalloc, particularly for short lengths, is so cheap it might as well be noise, which is cheaper than still fast array alloc for short length, and as the length passes the threshold to avoid excessive stack pressure, it becomes faster to retrieve pre-allocated array from a threadlocal bucket within shared array pool.
- neonsunset 2y agoIt might be, if you are writing low-level-ish code in C#, particularly one that uses its SIMD abstraction, this is something you do care about as it is relevant to extracting maximum instruction-level parallelism from modern deep and wide CPU cores. Pretty much the same knowledge that applies to C/C++/Rust applies to C# in such scenarios, save for swapping auto-vectorization consideration with the one for simpler usage of Vector128/256/512<T> (which, in turns, applies to the use of intrinsics in both the former and the latter).
- ykonstant 2y agoI know that SBCL can achieve C-like performance on some tasks, but I was not aware of a Scheme that did that; what are the implementations? Is any open source?
- widdershins 2y agoChez Scheme is probably the fastest well-maintained Scheme implementation. It's open source (MIT license). I would describe the performance as Go-like, rather than C-like, since it's garbage collected. But it's pretty darn good, especially for a dynamically typed language.
- ykonstant 2y agoThanks! This actually makes me wonder if the claims of SBCL C-like performance are out-of-date or exaggerated. I'd like to know if the SBCL compler produces SIMD-enabled, cache-friendly machine code and if so, by what mechanisms/annotations.