Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mbitsnbites
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
31.
▲
by
mbitsnbites
3y ago
BTW, do you have any blog, public repos or anything else?
32.
▲
by
mbitsnbites
3y ago
> I like that term. Do you have any suggested reading material from Alsup? He dwells in the comp.arch newsgroup, and is usually happy to answer questions. His take on RISC: https://groups.google.com/g/comp.arch/
33.
▲
by
mbitsnbites
3y ago
> ARM weren't even targeting the ultra low end, as they have a completely different -M ISA for that. That's the brilliance of it all, IMO. They didn't have to target the ultra low end, since they already had an ISA that
34.
▲
by
mbitsnbites
3y ago
Correction: Ok, the IBM z/Architecture line of CPU:s are clearly a different breed. In later generations they do use instruction cracking (i think that they were inspired by the x86 success), and insane pipelines: https://
35.
▲
by
mbitsnbites
3y ago
I see what your pointing at. I don't think that we'll fully agree on the nomenclature, but this kind of feels like the RISC vs CISC debate all over again. The reality is that the waters are muddied from the 1990's and onward.
36.
▲
by
mbitsnbites
3y ago
> I'm a huge fan of aarch64, it's a very well designed ISA I totally agree. I would go as far as to say that it's the best "general purpose" ISA available today. I am under the impression that the design was heav
37.
▲
by
mbitsnbites
3y ago
> Keep in mind that the average RISC ISA uses 5 bit registers IDs and uses three-arg form for most instructions, that's 15 bits gone. While AMD64 uses 4 bit register IDs and uses two-arg form for most instructions, which is only 8 b
38.
▲
by
mbitsnbites
3y ago
> The ZX Spectrum has a 3.5 MHz Z80 processor (1,000 times slower than current computers) Actually, it's much, much slower than that. The Z80 takes at least one clock cycle to add or subtract an 8-bit number (sometimes several clock
39.
▲
by
mbitsnbites
3y ago
BTW... Except for the indications from the die shots, one of the reasons that I don't think that uOPs can be as small as 32 bits is that studying fixed width ISAs and designing MRISC32 have made me appreciate the clever encoding tricks
40.
▲
by
mbitsnbites
3y ago
> I found annotated die shots of Zen 3 and Zen 4 Ooo thanks! Sure looks like strong evidence. > TBH, we have no idea how big the x86 tax is. No, and it gets even more uncertain when you consider different design targets. E.g. a 1000W
41.
▲
by
mbitsnbites
3y ago
Good points. I guess it's a debate of nomenclature, which lacks real importance (although it helps to reduce confusion). My point of view is mostly that, no, the x86 architecture certainly is not load-store, but internally modern x86
42.
▲
by
mbitsnbites
3y ago
> According to the label, that block contains both the uop cache AND the microcode ROM Yes, so it's hard to tell the exact size. We can only conclude that the uOP cache and the microcode ROM combined are about twice the size of the
43.
▲
by
mbitsnbites
3y ago
Thanks, I have read the Agner documents before. I will dig around some more and get updated. Anyway, I found this, regarding RMW (for Ice/Tiger Lake): > Most instructions with a memory operand are split into multiple μops at the all
44.
▲
by
mbitsnbites
3y ago
I'll have to read up on Agner's findings. My assumptions are largely based on annotated die shots, like this one of Rocket Lake (IIRC): https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_pr... If
45.
▲
by
mbitsnbites
3y ago
> The average A64 instruction does slightly less work than the average AMD64 instruction so the net result is that it makes sense to have slightly wider decoders I'm not sure that the difference is that big. A64 actually has quite p
46.
▲
by
mbitsnbites
3y ago
> Single-length instructions, for example. Even x86 instructions aren't much of a problem when you can afford to throw more pipeline stages at the problem I think that the problem is bigger than that. Sure, the branch predictor usua
47.
▲
by
mbitsnbites
3y ago
> 1) you could argue that's bad code generation, which could easily be fixed in the compiler, as `a1` could be used as the result of the `add` Yes, as I said. Please note that the code was generated by GCC (trunk) that has fairly ma
48.
▲
by
mbitsnbites
3y ago
I think that companies will continue to suggest a "big bang" that involves supporting integer instructions with three source operands and possibly dropping compressed instructions. Maybe it's out of selfishness (e.g. because
49.
▲
by
mbitsnbites
3y ago
Yes, but that's a big "if". When you start looking at it from the perspective "I don't need binary compatibility with microcontrollers", you realize that opcode space has been wasted on things that you don'
50.
▲
by
mbitsnbites
3y ago
I would say that all modern "general purpose" ISAs (and microarchitectures) are optimized for C (and C++). That's what most OS:es and high profile applications are written in (web browsers - and by extension electron apps, co
51.
▲
by
mbitsnbites
3y ago
In that case you don't need a canonical order. The instructions do not even have to be next to each other.
52.
▲
by
mbitsnbites
3y ago
Thanks a bunch for the link! I love how you fit a quake renderer onto such a small device (having a demo scene size coding history, I enjoy the challenge of a constrained target and I'd like to do something similar with the MRISC32 ISA
53.
▲
by
mbitsnbites
3y ago
I think you're onto something, but I'm not entirely sure that that's exactly how it's going to play out. First of all I don't think that "compatibility will reign". It's more like once the industry re
54.
▲
by
mbitsnbites
3y ago
Vircon32 looks very nice! So what did you do w.r.t. addressing modes? Do you have memory-memory operations? Is it a load/store architecture? I found the CPU specification ( https://github.com/vircon32/Vircon32Docume
55.
▲
by
mbitsnbites
3y ago
This was a very good read! Many spot-on observations. I actually started the design of the MRISC32 ISA before I knew about RISC-V (I even called it VRISC first, for "Vector-RISC", but had to give that name up for obvious reasons
56.
▲
by
mbitsnbites
3y ago
I have dodged most problems with speculation by executing the branch early, before the next instruction can enter the execute stage. That's one of the reasons why I have very simple branch conditions. E.g. I don't have the MIPS
57.
▲
by
mbitsnbites
3y ago
True, but for loads there is at least a theoretical way to avoid clobbers. For stores there is no way. Stores: 1-2 superfluous results Loads: 0-2 superfluous results
58.
▲
by
mbitsnbites
3y ago
It's a shortcoming in MRISC32 too. Unlike RISC-V, MRISC32 has a madd-instruction (integer multiply+add), which is helpful in multi-precision multiplication, but there is little help in terms of carry addition arithmetic. Together with
59.
▲
by
mbitsnbites
3y ago
Right now: assembly only. I don't really know compiler internals so it's an uphill battle.
60.
▲
by
mbitsnbites
3y ago
Checks out
More ›