4 ms·
>There are good reasons why 32-bit Arm failed to compete with Intel for performance, and why x86 has failed to displace Arm in low-power markets. The things tha
by hardware2win 3y ago
>There are good reasons why 32-bit Arm failed to compete with Intel for performance, and why x86 has failed to displace Arm in low-power markets. The things that you want to optimize for at different sizes are different.
What if x86 CPUs werent designed for low power? e.g due to focus on competitivness in perf focused markets
Saying that one isa is faster or more energy efficient is like saying that c++ syntax is faster than java syntax.
While there are lang features that enable stuff, then almost everything is up to the implementation - compiler, libraries, runtime and the programs code.
One letters arent faster than the other. ISA doesnt imply perf. characteristics of the end product.
Read this:
https://chipsandcheese.com/2021/07/13/arm-or-x86-isa-doesnt-matter/ https://chipsandcheese.com/2021/07/13/arm-or-x86-isa-doesnt-...
Even if you start talking about decoders, then:
>Another oft-repeated truism is that x86 has a significant ‘decode tax’ handicap. ARM uses fixed length instructions, while x86’s instructions vary in length. Because you have to determine the length of one instruction before knowing where the next begins, decoding x86 instructions in parallel is more difficult. This is a disadvantage for x86, yet it doesn’t really matter for high performance CPUs because in Jim Keller’s words:
- rbanffy 3y ago> Saying that one isa is faster or more energy efficient is like saying that c++ syntax is faster than java syntax Kind of. ISAs don't exist in a vacuum - for a given transistor budget, they'll force chip design choices that will drive power consumption and performance. Decoding instructions is one thing, but reordering them quickly and efficiently is more impactful for both power (if it can be done with fewer transistors) and performance (if it can be done better/faster so that more instructions from more instruction flows can be retired at the same time). I designed a beautiful ISA in college. It was a (mostly) stack machine with instructions designed to make a FORTH compiler extremely easy to implement. Unfortunately, it wouldn't be easy to evolve it past the point processors got faster than memory (I did not see that coming). It would, as originally designed, end up being unavoidably slow unless some fairly complicated caching were to be implemented. Another interesting example is the Intel 432 and its bit-aligned instructions. A lot of silicon that could be better used elsewhere was dedicated to fetching instructions. It was also slower to implement. On the x86 not being designed for efficiency, Intel has a whole line of CPUs designed for low-power environments. At some point, there was even a Motorola phone running Android on x86.
- snvzz 3y ago>On the x86 not being designed for efficiency, Intel has a whole line of CPUs designed for low-power environments. At some point, there was even a Motorola phone running Android on x86. Note it failed in the market, and was never competitive.
- rbanffy 3y agoIt failed on mobile phones. Atom and its descendants are in just about every cheap Windows tablet and small laptop, where being an x86 is advantageous.
- snvzz 3y agoWindows is (notably) not Android, and laptops/tablets do not run on the same power constraints a mobile phones does. Conversation went off the rails.
- GrumpySloth 3y ago> Saying that one isa is faster or more energy efficient is like saying that c++ syntax is faster than java syntax. Which may be true, if we’re talking about compilation speed, although in this case the reverse is true.
- jpcfl 3y ago> Saying that one isa is faster or more energy efficient is like saying that c++ syntax is faster than java syntax. I think that could be a valid statement. APIs can influence performance by constraining the implementation. For instance, the syntax for constructing an object in C++ will, generally speaking, always yield faster code than Java, because Java objects are almost always allocated on the heap, while C++ objects can be allocated on the stack. Compare: // C++ MyObj o{}; // vs. Java MyObj o = new MyObj(); Sure, it's possible to write a Java allocator/GC that will yield similar performance to the C++ code, but in general, that will practically never be the case. The syntax of the language has constrained the implementation so that Java will practically always be slower. Presumably, similar design choices in an ISA could have the same effect.
- hardware2win 3y ago>Sure, it's possible to write a Java allocator/GC that will yield similar performance to the C++ code, Didnt you just agree with me that it is dependent on the impl/end product? Because what would be the reasons in isa world to make it not desirable >The syntax of the language has constrained the implementation so that Java will practically always be slower. The most interesting question is: by how much? 1% 3%? 30?
- o11c 3y agoJava code is generally 2x slower than C++ code, unless you're operating entirely on primitive types. The JIT usually can't remove enough of the gratuitous memory accesses the language forces.
- hardware2win 3y agoYoure talking about java end to end, we are talking about just syntax.
- jpcfl 3y ago> Didnt you just agree with me that it is dependent on the impl/end product? No. I'm saying that it may be theoretically possible to tune performance in some cases, but not practical.