Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
eigenform
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
28 ms
·
61.
▲
by
eigenform
2y ago
Probably not, but I don't think anyone has talked about it explicitly. Otherwise, there are known examples of related-but-less-aggressive optimizations for resolving loads early. I'm pretty sure both AMD[^1] and Intel[^2] have h
62.
▲
by
eigenform
2y ago
Modern machines usually try to predict target addresses, not just the direction of conditional branches. You can implement it for unconditional calls and jumps too, even for direct/relative-addressed ones. That's pretty common now
63.
▲
Analyzing and Exploiting Branch Mispredictions in Microcode [pdf]
(arxiv.org)
3 points
by
eigenform
2y ago
|
0 comments
64.
▲
by
eigenform
2y ago
I dunno, you could imagine it happens [speculatively?] in parallel, at the cost of a read port and adder for each op that can be renamed in a single cycle: 1. Start PRF reads at dispatch/rename 2. In the next cycle, you have the result
65.
▲
by
eigenform
2y ago
> But I don't see how the bypass network could be bugged without causing the incorrect result. Maybe if they really rely on this kind of forwarding in many cases, it's not unreasonable to expect that latency can be generated by
66.
▲
by
eigenform
2y ago
Funny thing to double-check: are these encodings correctly specifying a 64-bit operand? Maybe everyone's compilers are subtly wrong D: edit: It looks like VEX.W is set in the encoding from the uops.info tests ¯\_(ツ)_/¯
67.
▲
by
eigenform
2y ago
Like, the scheduler is just waiting unnecessarily for both operands, despite the fact that 'count' has already been resolved at rename?
68.
▲
by
eigenform
2y ago
Oooh, I forgot about this. Only thing I can think of is something like: 1. Assume that some values can be treated internally as "physical register plus a small immediate" 2. Assume that, in some cases, a physical register is known
69.
▲
by
eigenform
2y ago
I wonder if this is a thing where the machine is trying to predict the actual value of the 'count' operand ...
70.
▲
by
eigenform
2y ago
Sorry, correcting myself here: it's cut across multiple cycles but not pipelined. Maybe I confused this with multiplication? If it were pipelined, you'd expect to be able to schedule DIV every cycle, but I don't think that&
71.
▲
by
eigenform
2y ago
Adding to this: the distinction is that an entire "instruction pipeline" can be [and often is ] decomposed into many different pipelined circuits. This article is specifically describing the fact that some execution units are pip
72.
▲
by
eigenform
2y ago
> Also IIRC there are still some non-pipelined units in Intel chips, like the division engine, which show latency numbers ~= to their execution time I don't think that's accurate. That latency exists because the execution uni
73.
▲
by
eigenform
2y ago
Yeah, I did a double-take when I read that too - but that does seem to be the case. From a different article [^1]: > "Throughout expert testimony, Arm has been asserting that all Arm-compliant CPUs are derivatives of the Arm instr
74.
▲
by
eigenform
2y ago
In the context of ARM machines, it's [historically] been the case that most of the devices are not servers (although that's slowly changing nowadays, which is nice to see!!)
75.
▲
by
eigenform
2y ago
It always seemed like [from ARM's point of view]: "oh, you're going to sell way more parts doing laptop SoCs with the license instead of servers... if we'd known that before, we would've negotiated a different licen
76.
▲
by
eigenform
2y ago
Now all we need is a RAM compiler to go along with the Factorio yosys plugin! (see https://mastodon.social/@thezoq2/112084897570820776 )
77.
▲
by
eigenform
2y ago
Presumably, when you have a relationship with ARM, you have access to things that make it somewhat less painful: - People who have been working with spec and technology for decades - People who have implemented ARM machines in fancy modern
78.
▲
by
eigenform
2y ago
You [as a designer] could probably add latency synthetically and still benefit from avoiding a physical register allocation (although I guess, that's only a workaround for leaking in the time domain). edit: Anyway, if your threat model
79.
▲
by
eigenform
2y ago
> I just don’t think Qualcomm seriously wants to invest in RISC-V unless ARM forces them to. That makes a lot of sense. RISC-V is really not at all close to being at parity with ARM. ARM has existed for a long time, and we are only now
80.
▲
by
eigenform
2y ago
That's true! - but still, I'm not so sure this is a reasonable comparison to make. If you're buying an Ampere workstation, it's reasonable to assume that your requirements are vastly different from someone buying a Mac
81.
▲
by
eigenform
2y ago
When you pay $3000 for an Apple Silicon machine, you're not paying for the same things as this. Totally different machine, power budget, application, use-case. edit: ie. sometimes you need a machine with 64+ cores running constantly at
82.
▲
Intel and AMD Form x86 Ecosystem Advisory Group
(amd.com)
11 points
by
eigenform
2y ago
|
0 comments
83.
▲
by
eigenform
2y ago
The difference is that, if we are solving a math problem together, you and I [explicitly or implicitly] can come to an agreement over the context and decide to restrict our use of language with certain rules. The utility behind our conversa
84.
▲
by
eigenform
2y ago
I just realized you were probably referring to the example given from the AnandTech article with `lea r64, [r64+imm8]`. Caveat is just that [presumably] the source and destination registers have to be matching (since `lea rax, [rax+imm]` is
85.
▲
by
eigenform
2y ago
> immediate addressing mode addition Well, except for the fact that you need to read from a register before adding the immediate displacement to it. You'd have to know the physical register and do the read very early (before renami
86.
▲
The Case of the Missing Increment
(computerenhance.com)
80 points
by
eigenform
2y ago
|
26 comments
87.
▲
by
eigenform
2y ago
The hardware itself is already very good at predicting the direction. I think the point about PGO here is that you need it to find infrequently-occuring biased-taken branches. When you encounter a biased-taken branch that isn't current
88.
▲
by
eigenform
2y ago
[Also not a geologist, but] apart from volcanism from subduction, there were also other landmasses being sheared off the subducting plate and accreted onto the western part of the continent. It was all probably higher before the many millio
89.
▲
by
eigenform
2y ago
I wonder if Sony having to adapt their DRM/platform security strategy into Intel-world would've introduced a lot of friction. This kind of thing is probably part of the motivation behind Intel splitting out a "Partner Securit
90.
▲
PDIP: Priority Directed Instruction Prefetching (ASPLOS '24) [pdf]
(cseweb.ucsd.edu)
3 points
by
eigenform
2y ago
|
0 comments
More ›