6 ms·
> It is a bit like assembly programming vs compilers. Sure clever programmers can beat most compilers when targeting simple processors, but when you level up wi
by pdhborges 14y ago
> It is a bit like assembly programming vs compilers. Sure clever programmers can beat most compilers when targeting simple processors, but when you level up with out-of-order executing, multiple instruction pairing, multiple level caches, NUMA, and so on, the compiler optimizer usually wins hands down.
I disagree. Real HP software have their computation kernels hand written (or generated) in assembly. As for relying on the smartness of the compiler to take into account all those factors I invite you to recall or read the story of the itanium architecture.
- pjmlp 14y agoFunny, I never saw hand written assembly while doing research at CERN. That seems pretty real HP to me.
- cmccabe 14y agoI work on Hadoop, and we do dip into assembly sometimes. It's very rare. One example is the CRC calculating code in HDFS. If you are working in a research context, it's not really worth writing assembly-- just getting it to work is more important. If you have to buy more hardware, then just do it. Traditional HPC is not really known for being very cost-sensitive. When you're working in a commerical context, performance starts to matter more. It's the same reason why a one-off hand-soldered electronic device isn't built to the same standards as an iPod. If you're only building one, don't waste time on polish.
- pjmlp 14y ago> I work on Hadoop, and we do dip into assembly sometimes. It's very rare. One example is the CRC calculating code in HDFS. Why aren't you using compiler vector operations for such case? On the project I used to work on, we were building the infrastructure to perform real time data analysis for the data coming straight out of the accelerator. We have written custom memory allocators, our own network stack and protocols, measured and optimized every operation in the years before the accelerator went live. The code is massively parallel in core and cluster distributed. No extra need for Assembly.
- cmccabe 14y agoIntel has a CRC instruction built in to the x86 instruction set on newer CPUs. It turns out that using this instruction is a substantial performance win. The problem with the "sufficiently smart compiler" argument, as always, is that the compiler may have a lot of optimizations, but it's still not an artificial intelligence. It can't tell that what you are trying to do with your set of operations is actually perform CRC, and there is hardware support for that. Check it out at: http://svn.apache.org/viewvc/hadoop/common/trunk/hadoop-common-project/hadoop-common/src/main/native/src/org/apache/hadoop/util/bulk_crc32.c http://svn.apache.org/viewvc/hadoop/common/trunk/hadoop-comm... We have written custom memory allocators, our own network stack and protocols, measured and optimized every operation in the years before the accelerator went live. I would actually argue that you don't need to do these things in most cases. SCTP and DCCP are alternatives to TCP/IP that have been in the Linux kernel for a while now. If you want to trash TCP/IP and go full custom, you can do so without writing a line of protocol code. However, a better approach is to use TCP/IP with tweaks like Fast Open (what Google uses internally.) You can then continue to buy commodity hardware. Similarly, memory allocators have been done to death. Just use tcmalloc or jemalloc rather than rolling yet another malloc(). The exception is if you want to create something like a slab allocator where you just hand out lots of mostly identically-sized buffers, or a database-like application where you manage your own writeback to disk.
- pjmlp 14y ago> The problem with the "sufficiently smart compiler" argument, as always, is that the compiler may have a lot of optimizations, but it's still not an artificial intelligence. It can't tell that what you are trying to do with your set of operations is actually perform CRC, and there is hardware support for that. This might be a special case, as you're taking advantage of a specific use case instruction, which will fail in portable code anyway. I have seen Assembly programmers put to shame in modern processors by C and C++ developers, just by making use of better data structures, algorithms and a special mix of compiler flags. Anyway thanks for the follow up, very interesting.
- cmccabe 14y ago