4 ms·
Why give up on performance? IA-64 (https://en.wikipedia.org/wiki/IA-64#Architecture https://en.wikipedia.org/wiki/IA-64#Architecture) already exists, with its
by lambda 9y ago
Why give up on performance?
IA-64 (https://en.wikipedia.org/wiki/IA-64#Architecture https://en.wikipedia.org/wiki/IA-64#Architecture) already exists, with its explicitly parallel instruction set, which leaves branch prediction and speculative execution up to the software. So you can still get many of the performance benefits, but it's under control of the software, giving a lot more flexibility for being able to mitigate or eliminate these kinds of issues.
I think that it was probably introduced ahead of its time, and targeted at the wrong markets, but it's kind of sad that with the Spectre vulnerabilities, we don't have any way of comprehensively addressing it in software without jumping through a lot of hoops and applying microcode updates.
- simias 9y agoIf you're not only concerned about bugs but also about people inserting backdoors into your hardware it's not enough to have access to the architecture design, you want to be able to publicly audit and review the entire manufacturing process (to make sure that the design you audited is actually the one etched on the silicon). I doubt Intel & friends would open their high-end fabs to public scrutiny so I think you'd have to use older technology available "in the open". For this reason alone I don't think a truly open processor could compete directly performance-wise with the current high-end CPUs regardless of architectural choices.
- deepnotderp 9y agoMost Itanium implementations had hardware branch predictors+predicates.
- Symmetry 9y agoTo get into trouble you have to be able to resolve one load and launch a new dependent one before the second load is quashed by the branch resolution. There exist in order pipelines that allow that, the Cortex-A8 for example, but they're pretty rare.
- deepnotderp 9y agoI'm just pointing out that Itanium had dynamic/hardware branch prediction and it's not necessarily immune to Spectre type exploits. As you mentioned, there are in-order scoreboarded pipelines that allow things like that, so dodging OoOE is not necessarily a perfect panacea.
- pjc50 9y agoI think this is overlooking how unpopular IA64 was. "Leaves branch prediction and speculative execution up to the software" means that in practice you need to either use the Intel compilers or hand-optimise your software to get the benefits, while simply recompiling your legacy software with GCC or LLVM ends up being disappointingly slow.
- lambda 9y agoOh, yeah, I'm aware of many of the reasons that IA64 failed. Intel could have solved this problem by contributing to GCC and LLVM instead of keeping their optimizations proprietary. Then they would also be more easily auditable as well.
- jasonwatkinspdx 9y ago> instead of keeping their optimizations proprietary There were no magic secret optimizations to release. It just straight up did not work. They had to add back dynamic branch prediction, and even then the load store latency was such trash that they had to put ginormous L3 caches on it to get even close to reasonable performance.
- lambda 9y agoThat's fair. I have never worked with them, so I don't know the precise details; but one of the big complaints I've heard was that compilers weren't ready for them yet, while these days there's now a larger field of open source compilers and has been a lot more research in parallelism.
- jasonwatkinspdx 9y ago> compilers weren't ready for them yet Yeah, compilers today are no better. We found the limits of statically scheduled parallelism pretty fast. On code that uses static scheduling, a modern OoO processor can easily duplicate what IA64 was capable of (and a pipelined loop using AVX will utterly smoke it), while being far better at all the stuff IA64 failed at.
- ryanlol 9y ago>targeted at the wrong markets What would have been the right ones?
- lambda 9y agoI'm not sure exactly what would have been the right markets, but I think that something where legacy compatibility is less important and being open to experimentation and new software designs is more important. So, things like consoles (going up against things like the Cell processor), HPC, or maybe even the embedded space (tablets, etc), rather than trying to use it for the enterprise market as Intel and HP did, where running legacy pre-compiled applications is a really important use case.
- pjmlp 9y agoGame developers are not at all open to experimentation, only adopting new languages or architectures when vendors force them to move forward. PS3 with Cell would have been a failure had Sony not eventually caved in, and released a PS3 SDK that made most of the work Sony was initially expecting devs to do, regarding low level programming. That work became PhyreEngine. https://en.wikipedia.org/wiki/PhyreEngine https://en.wikipedia.org/wiki/PhyreEngine http://develop.scee.net/research-technology/phyreengine/ http://develop.scee.net/research-technology/phyreengine/
- hajile 9y agoIA-64 was a dumpster fire. The halting problem is unsolvable. You can't schedule branch prediction absolutely. You have to have information about the current running program. Branch prediction, instruction re-ordering, and speculative execution are to hardware what a JIT is to software (roughly speaking). Unfortunately, from a security perspective, moving those things to software doesn't make the vulnerability go away. You can have the same vulnerabilities in your software implementation. From a performance perspective, we have no general solution to parallelizing a serial program. GPUs used to be VLIW. AMD switched from VLIW-5 to VLIW-4 because the average width was only ~3.5 (Nvidia had switched from VLIW to SIMD long before this). Today, Nvidia and AMD both use a MIMD threaded approach to execute on SIMD units. Later-generation Itanium chips wound up including branch predictors and speculative execution. From what I understand, under the hood, they were normal RISC-style processors (like all the x86 micro-arch are today). Just ignore the VLIW and run one set at a time serially with the ILP hardware optimizing as it goes. In today's programs, the programmer specifies data-level parallelism where possible (and if necessary, re-adjusts the code so the compiler heuristics recognize it as optimizable). The compiler then tries its best to detect the parallel data and use SIMD and organize instructions so that the parallelizable ones are closer together (so they fit in the CPU reorder buffer). When they hit the CPU, it examines the code as it runs to optimize speculative execution and uses the reorder buffer to make efficient use of its computation units (load, store, ALU, FPU, SIMD, etc). You move from explicit to implicit-ish (you know that putting similar instructions will optimize in all modern processors), but don't have the drawbacks of noop code bloat or having to compile different code when someone changes from VLIW-2 to VLIW-3 code.
- bogomipz 9y ago>"The halting problem is unsolvable" Can you elaborate on how the halting problem relates to the IA-64 architecture?
- MrMoenty 9y agoThe unsolvability of the halting problems means that it is impossible to have a general algorithm that makes non-trivial statements about the behavior of a given program. In the context of IA-64, that means that generally, the compiler will not be able to determine when it's save to use parallelism. We can only program it to use parallelism in a bunch of special cases for which we think it will be safe. Given that simias originally asked for a simpler platform that we can trust in more, this seems to be a pretty strong argument that IA-64 is not that architecture.
- deepnotderp 9y agoYou're giving up on a ton of performance by dropping OoOE. In particular, you end up losing the ability to avoid stalling on cache misses, and the memory wall becomes an increasingly gigantic problem even as Moore's Law progresses.
- nbsd4lyfe 9y agoI don't think you necessarily have to miss it on a threaded CPU, but that isn't as optimal.
- Narishma 9y agoHyper-threading helps with that.
- deepnotderp 9y agoNo, SMT techniques only help with throughput, not single threaded execution latency. Furthermore, if you have no OoO, then it's very likely that both threads will suffer a cache miss. Plus the overhead of SMT is very similar to just having two cores. Modern CPUs reuse much of their OoO circuitry for SMT, removing OoO means that SMT has greater relative overhead. (I can dig up the relevant papers if anyone's interested)
- Symmetry 9y agoIf you're looking for relatively high and deterministic performance maybe Project Denver, the descendant of the Transmeta's Efficeon, would be a better choice? Of course the binary translation has tons of potential for introducing bugs but since it has an exposed pipeline the translation software would have an easier time guaranteeing you wouldn't run afoul of Specter. NVidia's incarnation seems about as fast as an A57 which isn't great but isn't peanuts either. https://en.wikipedia.org/wiki/Transmeta https://en.wikipedia.org/wiki/Transmeta
- progman 9y ago> Why give up on performance? We don't need to give up on performance, even with low-performance RISC-V chips. What we need is a new architecture which takes advantage of low-speed multicore processors. If it is possible to produce thousand-core RISC-V chips [1] then any core could be assigned a single thread - which implies that there is no need for a complex multitasking OS. [1] Towards Thousand-Core RISC-V Shared Memory Systems (PDF) https://riscv.org/wp-content/uploads/2016/11/Wed1000-Thousand-core-RISC-V-Nguyen-MIT.pdf https://riscv.org/wp-content/uploads/2016/11/Wed1000-Thousan...