4 ms·
Why do we optimize our CPUs to fit our programming languages instead of desigining our languages around our CPUs?
by identity0 6y ago
Why do we optimize our CPUs to fit our programming languages instead of desigining our languages around our CPUs?
- WrtCdEvrydy 6y agoBecause languages are used for optimizing programmer output (these days) while CPUs are designed for general use compute. I think it would be a regression in technology to start having Intel Web-Compute Processors and Intel Gaming-Compute Processors but you know... that's just crazy me.
- pulse7 6y agoIBM has hardware that accelerates XML processing: https://en.wikipedia.org/wiki/IBM_WebSphere_DataPower_SOA_Appliances https://en.wikipedia.org/wiki/IBM_WebSphere_DataPower_SOA_Ap...
- infogulch 6y agoWow, if that's not proof that xml is a terrible format, I don't know what is.
- pjmlp 6y agoIt might not be optimal, but it certainly is better than the alternatives regarding machine processing, data validation, graphical tooling and comments.
- simion314 6y ago>Wow, if that's not proof that xml is a terrible format, A bad format for what? Do you mean that you assume that HTML,markdown would look better in something like JSON? I want to see a "proof" for documents you write in JSON by hand.
- pulse7 6y agoFor decades the CPUs have been optimized to run faster existing applications - regardless in which programming languages the apps are written. On the other hand you can find languages designed around our CPUs (like Rust with its "zero-cost abstractions").
- vbezhenar 6y agoIMO the last language designed around CPU was C designed around CPUs of that time. Modern CPU have registers, caches, SIMD instructions which are tremendously important for performance. Yet I don't know no language which would explicitly provide any control over those things. You can only hope that your loop will be understood by compiler and optimized to proper SIMD instructions. And any change might break that optimization (including changes in compiler).
- dragontamer 6y ago> Yet I don't know no language which would explicitly provide any control over those things. CUDA provides your explicit cache control through "__shared__" memory. C commonly has low-level attributes. In particular, you have the "register" keyword in C if you really want to go that route, but compilers have a nice greedy algorithm that is damn near optimal, so its best to let the compiler perform its own analysis for register allocation. SIMD instructions are easily provided through intrinsics: https://software.intel.com/sites/landingpage/IntrinsicsGuide/ https://software.intel.com/sites/landingpage/IntrinsicsGuide... ------- There's plenty of low level languages that provide the control you desire. The main reason for __shared__ memory is that compilers can't optimize memory across threads yet, so we still require the programmer to tile memory manually and think about the optimizations across the many, many threads a CUDA multiprocessor supports. If you want to explicitly control the cache on x86 systems, use the _mm_prefetch() intrinsic. In practice though, it is pretty rare for prefetching to be useful on modern systems. (Auto-prefetchers will optimize most sequential traversals, and those that fail prefetching can still be out-of-order executed at decent speeds)
- gleenn 6y agoThe reason we have programming languages at all is to make the process of using the CPU easier. We don't code in assembly because it's easier to do things with higher level languages. If you can make a language both designed to run around CPU hardware well and also be easy for a human to use, then great. If not, then it's trade-offs. If I can make a higher level language like Java run better on hardware, then isn't that great? If it makes the CPU slower because the Javaness of it, then I guess that sucks, but it isn't necessarily the end goal.
- pjmlp 6y agoBecause we have been doing it for decades, mainframes have always done like that, Xerox PARC workstations used microcoded CPUs, since UNIX and C got widespread, CPU vendors optimize for C including memory tagging as hardware mitigation, and NVidia nowadays designs their cards for C++ and Tensorflow workloads.
- eru 6y agoWe have been doing both for a long time.
- SomeoneFromCA 6y agoC was designed around PDP-11.
- schoeberl 6y agoAnd now we design processors to run C fast.
- jecel 6y agoThough RISCs are considered language agnostic, they tend to be optimized for C. One side effect of that is the elimination of rotation instructions which C doesn't have (though GCC can generate them as a special case). RISCs also don't help with lexically scoped variables which languages like Pascal have but C doesn't. A processor that was specifically designed for C was called CRISP (and was briefly sold as the AT&T Hobbit): https://en.wikipedia.org/wiki/AT%26T_Hobbit https://en.wikipedia.org/wiki/AT%26T_Hobbit
- exabrial 6y agoI'd much prefer we design processors around languages, because from what I understand, x86 isn't really optimized for anything. x86 itself has a lot of glut. Most x86 processors don't directly implement x86 in hardware but execute instructions in microcode, much like JVM and Java bytecode.
- dragontamer 6y agoSerious question: Why not both? OpenCL / CUDA are good examples of languages designed around GPUs. C itself was based off of the DEC-PDP11. Optimizing the other way: Many modern processors are based on executing C faster and faster... but other languages have begun to be targeted by CPUs. Not only this JOP, but ARM famously has the "convert to javascript floating point" instruction, to optimize our phone's handling of Javascript and HTML pages. There's lots of engineers. Why not have all the engineers try to make everything faster? What benefit is there of doing things only one at a time as you propose?