8 ms·
If you watch the Belt talk on our site, and know how a modern OOO machine works on the inside, then you will recognize that the Belt is a forwarding network, so
by igodard 12y ago
If you watch the Belt talk on our site, and know how a modern OOO machine works on the inside, then you will recognize that the Belt is a forwarding network, sometimes also called a bypass. There is no RAM, "S" or otherwise, no general registers, and no ports.
Bypasses are nothing new; what is novel is how we are able to handle triple the number of data paths than other machines, the fact that the bypass is exposed to the program rather than being hidden behind the register metaphor, and that the program model is a single-assignment FIFO. See those talks for more.
There is no better way to speed up registers and SRAM than by having none at all.
- hahainternet 12y agoThanks for commenting Ivan. I'm excited for the future and hope to see more material produced by you. The original belt video absolutely blew my mind and I feel that Mill has the opportunity to really revolutionise the entire industry.
- thesz 12y agoOkay. Your model looks like TTA - a machine that is built on bypasses. They are not new and they can be used to create very efficient chips (in terms of operations/watt) for some fixed functions (precisely, FFT of 2^N). But they are 1) not fast in terms of raw performance for general purpose tasks, 2) not fast in terms of operating frequency and most important 3) prone to stall when present with non-deterministic delays like access to RAM. You can add whatever functional units you like to TTA design, including content-addressable memory in disguise as FIFO. TTA design with such device will be identical to what you've described above. I won't think you will improve performance very much with this trick. PS "not fast for general purpose tasks" - in some benchmarks TTA architectures executed gcc 10+ times slower than general purpose CPU with same frequency. "not fast in operating frequency" - TTA requires crossbar, which is slow in 2D. You cannot make it fast. "prone to stall" - you have to stop complete pipeline for a cache miss, otherwise you'll have divergence in execution.