Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
BeeOnRope
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
61.
▲
by
BeeOnRope
2y ago
Yes I don't really get it. For non-loop branches if the compiler has knows enough to insert a branch hint it would also already arrange the generated code so that was the fall through cases, which is already much more efficient when pr
62.
▲
by
BeeOnRope
2y ago
What's special about the E core's L2 cache such that it gets on-chip regulated voltage?
63.
▲
by
BeeOnRope
2y ago
Sounds a bit like the jcc erratum?
64.
▲
by
BeeOnRope
2y ago
Modern compilers are not doing much searching in general. It's mostly apply some feed-forward heuristic to determine whether to apply a transformation or not. I think a slower, search based compiler could have a lot of potential for th
65.
▲
by
BeeOnRope
2y ago
I'm not following: as long as you are introducing a new, incompatible instruction for leading zero counting, you'd definitely choose LZCNT over BSR as LZCNT has definitely won in retrospect over BSR as the primitive for this use c
66.
▲
by
BeeOnRope
2y ago
Probably, if the uops come from the uop cache you get the fast speed since the prefix and any decoding stalls don't have any impact in that case (that mess is effectively erased in the uop cache), but if it needs to be decoded you get
67.
▲
by
BeeOnRope
2y ago
Yes, files can be sparse but the actual disk usage information is also returned by these stat-family calls, so there is no special cost to handling sparse files.
68.
▲
by
BeeOnRope
2y ago
This seems difficult since I'm not aware of any way to get approximate file sizes, at least with the usual FS-agnostic system calls: to get any size info you are pretty much calling something in the `stat` family and at that point you
69.
▲
by
BeeOnRope
2y ago
The rules and mechanisms for SMC detection are essentially the same in both modes as far as I am aware. Both Intel and AMD implement SMC detection that is a bit stronger than required by the specification as well.
70.
▲
by
BeeOnRope
2y ago
> of the high-performance programs that is possible in 512-bit AVX-512 due to the equality between register size and cache line size, so the consumer Intel CPUs will remain a worse target for the implementation of high-performance algori
71.
▲
by
BeeOnRope
2y ago
It seems like that would be a likely common behavior for the FTL, but other options are possible (e.g., reading the old blocks) and it wasn't guaranteed by the spec, which is why they added this NVMe flag (so-called "DZAT") s
72.
▲
by
BeeOnRope
2y ago
Not guaranteed by default for NVMe drives. There's an NVMe feature bit for "Read Zero After TRIM" which if set for a drive guarantees this behavior but many drives of interest (2024) do not set this.
73.
▲
by
BeeOnRope
2y ago
Yes.
74.
▲
by
BeeOnRope
2y ago
libc++ string is smaller with a higher SSO capacity which in many scenarios can overwhelm the code generation. So it's hard to draw an absolute conclusions.
75.
▲
by
BeeOnRope
3y ago
That's not how it works on x86 as far as I know. The atomic ops are simply performed by the ALU, against the L1 cache when the instruction is about to retire. Atomicity is guaranteed by not allowing the line to be stolen by another cor
76.
▲
by
BeeOnRope
3y ago
> Doesn't explain why there's no -fshadow-stack-only option to pass in. I thought you were asking about the design of the hardware: it's designed that way because compatibility means that the vast majority of people want s
77.
▲
by
BeeOnRope
3y ago
ABI compatibility.
78.
▲
by
BeeOnRope
3y ago
Which is the slow unwinding path? The one from libbfd?
79.
▲
by
BeeOnRope
3y ago
This kind of memory order speculatiom is basically required on x86 since the strong semantics would otherwise prevent many useful reorderings (especially L-L which is absolutely critical). The basic way it works is that pretty much any reor
80.
▲
by
BeeOnRope
3y ago
In EC2, most of the "storage optimized" instances (which have the largest/fastest SSDs) generally have more advertised network throughput than SSD throughput, by a factor usually in the range of 1 to 2 (though it depends on e
81.
▲
by
BeeOnRope
3y ago
> But that doesn't prevent issues due to switching between customers on the same physical core, no? Yes they are explicit that customers may be time-shared on a physical core ("burstable" instances don't really make s
82.
▲
by
BeeOnRope
3y ago
Granted but the vendors accept these are serious problems given that they are immediately patched and the mitigation all enabled by default even at significant performance cost (most chip generations are down double digit perf % based versu
83.
▲
by
BeeOnRope
3y ago
Shouldn't the functions of "development" and "QA" both reside under the umbrella of the chip-maker though? In fact, chip-makers famously invest an insane amount of money into "QA" (aka "validation&quo
84.
▲
by
BeeOnRope
3y ago
That's true, but it leads the odd assumption that the vendor managed to fix N side-channel attacks before release but 0 thereafter, while random individuals fixed M thereafter over a period of years with N >> M. This seems to be
85.
▲
by
BeeOnRope
3y ago
I dabble in this space (hardware reverse-engineering) and write software for a living and in my opinion the gaps are huge. I should disclose have been paid by a chip-maker for a blog post that I wrote which "disclosed" an optimiza
86.
▲
by
BeeOnRope
3y ago
Nevermind, AWS explicitly documents that all instance types, including burstable, never co-locate different tenants on the same physical core at the same time: https://docs.aws.amazon.com/whitepapers/latest/securit
87.
▲
by
BeeOnRope
3y ago
I'm not directly in charge of hiring QA, no! I think this sort of excuses the initial blindness to Spectre style attacks in the first place, but once the basic pattern was clear it doesn't excuse the subsequent lack of discovering
88.
▲
by
BeeOnRope
3y ago
> A fundamental problem is that the attack surface is so, so huge. Even if their security researchers are doing blue-sky research on both very small and very broad areas of processor functionality, they're going to miss a lot. Sure.
89.
▲
by
BeeOnRope
3y ago
I think the comparison between CPU and software exploits holds at a very high level, but in the case of software the gap between internal and external researches seems lower. Much software is open-source, in which case the play field is alm
90.
▲
by
BeeOnRope
3y ago
> Also you’re dealing with a company that has been running to stand still for a long time I'm not just talking about Intel, but also Arm and AMD. As far as I know none of these has obviously been making proactive Spectre fixes.
More ›