5 ms·
First time I'm hearing about Tachyum. Looks very promising, although it does sound too good to be true, especially the cross-ISA support. Interesting bit I fou
by fathyb 4y ago
First time I'm hearing about Tachyum. Looks very promising, although it does sound too good to be true, especially the cross-ISA support.
Interesting bit I found researching them: https://www.nextplatform.com/2020/04/02/tachyum-starts-from-scratch-to-etch-a-universal-processor/ https://www.nextplatform.com/2020/04/02/tachyum-starts-from-...
The processor pipeline has its out of order execution handled by the compiler, not by hardware, so there is some debate about whether this is an in order or out of order processor. Danilak says that instruction parallelism in the Prodigy chip is extracted using poison bits, which was popular with the Itanium chip which this core resembles in some ways and which are also used in Nvidia GPUs. The Prodigy instruction set as 32 integer registers at 64-bits and 32 vector registers that can be 256 bits or 512 bits wide, plus seven vector mask registers. The explicit parallelism (again, echoes of Itanium) is extracted by the compiler and instructions are bundled up in sizes of 3, 8, 12, or 16 bytes.
- rayiner 4y agoSo it’s a big VLIW chip. Maybe interesting for the domain, but the crazy specs make sense in that context. With deep OOO designs hitting almost 5 ghz on TSMC 5 nm, it’s not surprising to see an in order VLIW design hit 5.7. Impressive effort by new company if it pans out though. It would be good to see a renaissance in high end chips and architectures like the mid 1990s.
- Symmetry 4y agoYeah. I wonder if its a classic VLIW or if the pipeline is exposed and/or skewed. exposed: The results of operations that take more than one clock cycle don't necessarily appear at their destination the cycle after the instruction is executed. Think branch delay slots but potentially for multiplies and loads too. skewed: Loads, processing, and stores can happen on subsequent clock ticks so simple loops don't necessarily need prologues and epilogues. A compiler can handle either of these for a particular CPU pretty easily but they tend to eliminate binary compatibility. Code morphing as in the Transmeta lineage like to use both of these. Other VLIWs with barrel multithreading, like the Hexagons DSPs in Snapdragon SOICs, don't need them. EDIT: The Mill guys have a video on how this works on their system. They've got their own names for things for some reason but they do a good job of explaining how this all works: https://millcomputing.com/docs/execution/ https://millcomputing.com/docs/execution/
- pclmulqdq 4y agoIn my time as an FPGA developer, I wrote a lot of tiny processing cores that were exposed and skewed VLIW designs for specific applications. If the code isn't changing much, they work extremely well in a very small silicon area and power footprint. However, they are awful to program. In supercomputing, this sort of thing sometimes appears and works well (like the Pezy-SC2 chips that recently came out). However, they usually fail on general purpose computing tasks.
- Symmetry 4y agoYeah. Normally sophisticated software scheduling works great on DSP-like tasks where the memory access patterns are very predictable but suffer a lot where you tend to have unexpected cache misses that a deep OoO system could just paper over.
- londons_explore 4y agoI think the benefits of both could be had with "software assisted branch prediction/caching". Ie. you transpile your x86 code to native code for your VLIW machine. Then you run that code for a few hundred clock cycles till bam "cache not ready on time exception" which fires when you try to execute an instruction that is expecting to read memory from the cache, but the cache isn't yet populated with that value. Then you re-run your transpiler which will produce new code which either does a better job of reading the data into the cache ahead of time, or issues different instructions which take more time and read data direct from RAM. Remember a software transpiler sometimes has more information than a typical deep OoO CPU, because it can use a lot more memory for state (eg. remembering that a particular branch or memory access won't be cached), and it can even persist state across system reboots. It can also do far more expensive optimizations and save the results, something a deep OoO CPU can't do because all the optimizations need to be doable in hardware.
- freemint 4y agoAll i can think is online polyhedral optimization using the unused matrix instructions. I recoil in horror.
- deleted 4y ago[deleted]
- Symmetry 4y agoI wonder how they manage that? It's straightforward to do if a page fault represents an error that halts the execution of a thread but not if you need to be able to keep going after the OS reads in the data from the hard drive. That is, if you speculatively load values that aren't going to be used in a loop then the fact that you had poisoned values sitting around in some registers doesn't matter. But if that data "should" have been there but wasn't and you store or branch based on it then you've got to roll back to the load, page in the data, then continue from the checkpoint. And if you're able to do all of that why not just go full out of order?
- rayiner 4y agoThe pipeline state can be checkpointed with a fixed amount of logic gates that’s proportional to the amount of pipeline state. Out of order execution requires tracking dependencies in structures that scale superlinearly with the size of the out of order window.
- Symmetry 4y agoIt isn't bound by the pipeline length. Even without pipelining at all, with poisoning you can do a faulting load into a register and then just mark the value in the register as poisoned. As you perform operations with the value any results are also poisoned. Then maybe dozens or hundreds of cycles later when you store or branch on one of the poisoned values the fault occurs. So pipeline length really doesn't have anything to do with it. On Itanium you essentially had to double check your values for poison before operating on them in a lot of cases and that caused big performance penalties on conventional code. Other systems restrict the shenanigans you can get up to with your MMU and solve the problem that way but that means you can't run a traditional OS with mmaped files and memory paging and such. I'm not sure what the Tachyum people are doing or if they've got some clever idea to get around all of this.
- avianes 4y agoI'm not sure if this answers your question but VLIW processors feature speculative loads. "speculative load" can be seen as data prefetch to a register, but with a extra instruction to put before reading/using the register. This instruction check that the memory content has been received, otherwise the instruction stall the pipeline. It allows to reorder a load before a conditional branch. To determine whether the load has been received, there is an additional structure in the microarchitecture that tracks speculative loads in-flight. When a context switch occurs, this data-structure is overwritten, so there are additional mechanisms to replay the load in such cases.
- fweimer 4y agoRegarding the “too good to be true” part, I wouldn't take them seriously based on what they have published (and more crucially what they have withheld, like an ISA manual or sources of a GNU toolchain port), except they managed to hire a very senior GCC developer. I trust that he did his due diligence before joining them, so I must assume what they are building is real. Did they make any performance claims about the cross-ISA support? It might just be QEMU port with qemu-user and TCG.
- masklinn 4y ago> except they managed to hire a very senior GCC developer. I trust that he did his due diligence before joining them, so I must assume what they are building is real. "Lots of money" and interesting challenge can be the due diligence, it's not like the product has to be realistic to get paid (it helps in the long run, but in the short run VC money pays the bills). Hell, they might even believe in the product. Linus spent 6 years at Transmeta.
- pinewurst 4y agoI think there's a difference though between Transmeta not making a competitive product from their story and Tachyum seeming an out-and-out scam.
- masklinn 4y ago> Tachyum seeming an out-and-out scam. But if you actually look at the thing, it does not actually seem like an out-and-out scam, a 5.7GHz VLIW design is not even remotely in the realm of the impossible. And rosy promises which end up crashing and burning are exactly what transmeta achieved.
- foobiekr 4y agoTransmeta managed to hire Linus Torvalds. Hiring is mostly - not entirely - about money.