5 ms·
Thanks david for providing that reference. I missed the discussion the first time. Perhaps you, or someone here, would know: has anyone asked and tried to anal
by throwaway000002 10y ago
Thanks david for providing that reference. I missed the discussion the first time.
Perhaps you, or someone here, would know: has anyone asked and tried to analytically answer what primitive operations an ALU ought to have? I mean, everything could be coded as a look-up-table, but given code "on average", what should be available in hardware.
It's a strange question that requires you pose it properly to even begin to answer it. For example, in the kind of stuff I do, popcnt and other "set-like" bit operations, i.e. most-significant set bit, are important enough that I just wish they were built in.
Mine you, I'm doing high-performance data processing, so perhaps it's not part of the scope of "average" code, but I disagree.
- dewster 10y ago"...has anyone asked and tried to analytically answer what primitive operations an ALU ought to have?" Very good question. I think this is where the "art" of computer design starts, and it generally gets short shrift. For instance, the ARM lacks a leading zero count instruction, which is a pretty wild omission because it has tons of uses (particularly for floating point) and is fairly expensive to implement in software using other primitives. I know it's not scientific, but I learned what to put in the Hive ALU by programming various algorithms I figured I would need at some point. I just developed a bunch of floating point subroutines for Hive (cos, sin, sqrt, 2^x, log2, 1/x, etc., not in the paper yet) and, short of adding a floating point pipeline, could only really justify adding an opcode that returns +1 / -1 based on the sign bit (for use in storing sign and doing absolute value). My seat of the pants rule is that if it has broad applicability, saves time and code space, and isn't too costly in terms of hardware / speed, then I put it in. Otherwise I leave it out. IMO, more of computer engineering should focus on what to leave out.
- stephencanon 10y ago> For instance, the ARM lacks a leading zero count instruction Huh? ARM has this instruction (since v5, IIRC), it's spelled CLZ.
- dewster 10y agoI've read that ARM Cortex-M0 and M0+ cores do not support it. Why on earth isn't it in all ARM cores?
- stephencanon 10y agoM0 is an extraordinarily stripped-down core. While I agree that it would be nice to have CLZ, if you are really missing it on an M0, you'd probably be better off with M3 or M4f.
- throwaway000002 10y agoI had a glance at your Hive paper. Seems like very interesting work. I don't have all the know-how, but if I had the resources, I'd like to build a system where I push compute to a node that essentially is a M.2 SSD glued to a custom core glued to 10G networking. Wire it all up in a Clos network. Use IPv6 and treat it as the address space, throw in a few routing tricks, and call the whole thing a computer. One day I'll build it. Hopefully I can get to that scale at some point.
- dewster 10y agoI'm not sure why more processors aren't tightly integrated into memory (and vice versa). Physically separating processor and memory with slow package pins means caches and all the real-time headaches and complexity and die size / cost issues they bring with them. When I assemble a new PC I almost never increase the memory down the road, so it might as well be on the processor die.
- noir_lord 10y agoThey probably will be, Integration is the future once all the other low (and not so low) fruit has been collected.
- Animats 10y ago"Has anyone asked and tried to analytically answer what primitive operations an ALU ought to have?" Do you want to go fast, or do you want to get the gate count down? Intel does code measurement to decide which instructions have to go fast, and which don't. Typically they run Windows workloads for this. It's not an analysis from first principles. The RISC crowd put a lot of effort into trying to reduce the instruction set. They tended to end up with simple instructions that could be hard-wired, and lots of registers. Many early microprocessors were microprogrammed to get the gate count down, and slow because of it. RISC was a reaction to that. Early RISC thinking focused on not only making all instructions the same length, but making them all take the same time - one instruction per clock. Then superscalar came to microprocessors, with the Pentium Pro, and ordinary CPUs started getting more than one instruction per clock. (Before the Pentium Pro, only supercomputers had that kind of hardware.) This killed the basic advantage of RISC. Most of the little Forth CPUs are one instruction per clock. Most Forth CPUs can hit the data stack, the call stack, and the main memory, which are all separate, on each cycle.
- gpderetta 10y agoThe Pentium was the first superscalar x86 cpu, at last from Intel. I'm pretty sure there were plenty of superscalar riscs around that time. Pentium pro was the first intel Out Of Order Execution x86. It was a great cpu (it's architecture was used for the next 10 years), but again, there were plenty of workstation and server class ooo RISC at that time. Edit: missing ooo