Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
BeeOnRope
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
BeeOnRope
2y ago
Really interesting. Normal uops don't work like that, they are always pipelined, so a p06 op with 3-cycle latency would always be 3/0.5, not 3/1. So the 1-throughput strikes me as a renamer limit, not an execution limit. I.e.
32.
▲
by
BeeOnRope
2y ago
How did you test the throughout? Seems quite weird. Does it go to both ports still?
33.
▲
by
BeeOnRope
2y ago
https://news.ycombinator.com/item?id=42582623
34.
▲
by
BeeOnRope
2y ago
The slow state is when rcx has a special "hidden immediate" attached it, from the `mov rcx, 1` or `add rcx, 1`. If the last option on the register doesn't fit that narrow template, it is fast, so loading from memory in any wa
35.
▲
by
BeeOnRope
2y ago
Yeah I am pretty sure that the renamer adds the tracked immediate(s) to the emitted uop, but that's not inconsistent with tracking "reg offset pairs" is it?
36.
▲
by
BeeOnRope
2y ago
It's true, I was considering only the original mov rax, 1 case, which I'm pretty sure compilers don't generate: it's just a useless encoding of mov eax, 1. Given that 64-bit immediate math also causes this though, compil
37.
▲
by
BeeOnRope
2y ago
Does this also occur with other 3-argument instructions like ANDN?
38.
▲
by
BeeOnRope
2y ago
To rule out alignment you should adding padding to one so the two variations have the same alignment in their long run of SHLX (I don't actually think it's alignment related though).
39.
▲
by
BeeOnRope
2y ago
Worth noting that whether intentional or not, this would be easy to miss and unlikely to move benchmark numbers since compilers won't generate instructions like this: they would use the eax form which is 1 byte shorter and functionally
40.
▲
by
BeeOnRope
2y ago
Division is complicated by the fact that it is a complex micro-coded operation with many component micro-operations. Many or all of those micro-operations may in fact be pipelined (e.g., 3/1 lat/itput) , but the overall effect of
41.
▲
by
BeeOnRope
2y ago
We are interested in the software visible performance effects of pipelining. For small benchmarks that don't miss in the predictors or icache, this mostly means execution pipelining. That's the type of pipelining the article is di
42.
▲
by
BeeOnRope
2y ago
Yes. In the past new HW has been made available to the uops.info authors in order to run their benchmark suite and publish new numbers: I'm not sure if that just hasn't happened for the new stuff, or if they are not interested in
43.
▲
by
BeeOnRope
2y ago
Division, gather are muti-cycle instructions which have typically had little or no pipelining on x86.
44.
▲
by
BeeOnRope
2y ago
Yes, it applies to different operations. E.g. you could interleave two or three different operations with 3 cycle latency and 1 cycle inv throughput on the same port and get 1 cycle inv throughput in aggregate for all of them. There is no r
45.
▲
by
BeeOnRope
2y ago
I don't think anyone is talking about "fetch, decode, operate, retire" pipelining (though that is certainly called pipelinig): only pipelining within the execution of a instruction that takes multiple cycles just to execute (
46.
▲
by
BeeOnRope
2y ago
That's true, but another part of the tables show how many "ports" the operation can be executed on, which is enough information to concluded an operation is pipelined. For example, for many years Intel chips had a multiplier
47.
▲
by
BeeOnRope
2y ago
So does the TV have to scan all RF channels at startup to build a virtual->RF channel map?
48.
▲
by
BeeOnRope
2y ago
> There are also some EC2 instance classes where upgrading instance types in the same "size" are more expensive An increase in price has been the rule rather than the exception for recent upgrades for vanilla instance types, e.
49.
▲
by
BeeOnRope
2y ago
Yes it the same for Intel. One quantitative difference is that tasks assigned to the E cores may run at a sustained frequency much lower than the maximum (3.7x here) while on Intel any sustained load generally results in frequency scaling u
50.
▲
by
BeeOnRope
2y ago
They won't be mispredicted nor take predictor resources since the default prediction is "not taken" and these branches are never taken (except perhaps once immediately before an OOB crash if that occurs). So they are free in
51.
▲
by
BeeOnRope
2y ago
> When that trick can't be used, I think the most efficient method would be to clamp the top of the address so that the max would land on a single guard page. If you are already doing a cmp + cmov, wouldn't you be better off ju
52.
▲
by
BeeOnRope
2y ago
> Can you be more specific about "all execution after a mispredict is thrown away". Are you saying even non-dependant instructions? To clarify, the misprediction happens at some point in the linear instruction stream, and all i
53.
▲
by
BeeOnRope
2y ago
Nice article! Very cool to see both the "independent" and "serially dependent" cases addressed. Microbenchmarks still have lots of ways of giving the wrong answer, but looking at both these cases exposes one of the big v
54.
▲
by
BeeOnRope
2y ago
I guess we are probably in violent agreement on most of this.
55.
▲
by
BeeOnRope
2y ago
> Spectre mitigations don't change that, ... Yes, exactly. To the first order I think Spectre didn't really change the performance of existing userspace-only code. What slowed down was system calls, kernel code and some things
56.
▲
by
BeeOnRope
2y ago
No, just a consequence of how mispredicts work: all execution after a mispredict is thrown away: though some traces remain in the cache, which can be very important for performance (and also, of course, Spectre).
57.
▲
by
BeeOnRope
2y ago
To answer a question implied in the article, per-lookup timing with rdtscp hurts the hash more than the btree for the same reason the hash is hurt by the data-depending chaining: rdtscp is an execution barrier which prevents successive look
58.
▲
by
BeeOnRope
2y ago
> If we define 'processed' as fully committed, then messages ca be processed exactly once. If it can be processed once, it can be delivered exactly once if you accept the simplest definition of that word. I don't see how.
59.
▲
by
BeeOnRope
2y ago
In a sense it's just kicking the can down the road: how you ensure the consumer reads the message only once? It may fail at some point after it has read the message but updated some state to reflect that. So it's not just a pointl
60.
▲
by
BeeOnRope
2y ago
That's exactly what it does by default. Only if you pass --and-rebase does it actually do the autosquash rebase for you.
More ›