Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
BeeOnRope
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
91.
▲
by
BeeOnRope
3y ago
I considered this, but we have pretty good evidence that the chipmakers have not been busily secretly patching Spectre attacks: 1) Microcode updates are visible and Spectre fixes are hard to hide: most have performance impacts and most requ
92.
▲
by
BeeOnRope
3y ago
What I find odd is that after the initial Spectre attacks, there have been a long string of these attacks discovered by outside researchers and then patched by the chipmakers. In principle it seems like the chipmakers should hold all the ca
93.
▲
by
BeeOnRope
3y ago
They are definitely time-sliced among tenants and very possibly two tenants may run at the same time on two hardware threads on the same core: but you could have a viable burstable instance with time-slicing alone.
94.
▲
by
BeeOnRope
3y ago
> These are not the same thing. Afaik, most “vCPU” are hyperthreads, not physical cores. OP didn't say otherwise. They are saying that public clouds do not let work from different tenants run on the same physical core (on differe
95.
▲
by
BeeOnRope
3y ago
Yes because of the "it's the cold of path" aspect. That said, CAS is often considerably slower than an atomic increment replacement even if the usual sources report equal throughout because the dependency chains to both the o
96.
▲
by
BeeOnRope
3y ago
4 might be useless compared to 3 from a "required quorum" PoV but if you want to stay up with 1 failure and the load can't be handled on 3 servers but 4 is OK, it is optimal isn't it? I.e. selection of ideal server coun
97.
▲
by
BeeOnRope
3y ago
Can you use the kernel polling mode w/o running as root?
98.
▲
by
BeeOnRope
3y ago
We weren't necessarily talking about that at all but whether data "lost" because a client crashed before it received acknowledgement of a durable write from the server is somehow the same as losing data that has been acknowle
99.
▲
by
BeeOnRope
3y ago
This is not data loss sense we talk about for Kafka or other queues, however, since the messages have not been acked: the state of unacked messages is completely unknown and no guarantees are made about them.
100.
▲
by
BeeOnRope
3y ago
> Batching should be done on the client side anyway, as most packages already do by default. If you are worried about too many fsyncs degrading performance, batch harder on your clients. It's the better way to batch anyway. This is
101.
▲
by
BeeOnRope
3y ago
x86-64 is not even a particularly compact encoding even compared to contemporary fixed length encodings. The inherent advantage of variable length encoding is largely cancelled out by wasted encoding space for legacy cruft. Aarch64 is rough
102.
▲
by
BeeOnRope
3y ago
Having a wide backend which cannot be fully utilized is common in non-x86 chips as well. The backend units are specialized while the front end less so, so to sustain the front-end bandwidth for many instruction mixed you need a wider backen
103.
▲
by
BeeOnRope
4y ago
They are trying to solve this problem without rewriting the branch.
104.
▲
by
BeeOnRope
4y ago
I think side effect is a red-herring here: it's idiomatic to rely on short circuiting when the RHS relies on the LHS being true to be safe to call at all. In the above example, we check that the pointer is non-null before calling it.
105.
▲
by
BeeOnRope
4y ago
> Relying on short circuiting to avoid side effects or improving performance meaningfully is pretty shaky. Says who? IME relying on short circuiting is useful and idiomatic in C and C++ code, e.g.: if (callback && callback-
106.
▲
by
BeeOnRope
4y ago
But radix sort doesn't write randomly to the output array like that. Yes, each element ends up in the right place, but through a series of steps which each have better locality of reference than random writes. In this way, radix sort i
107.
▲
by
BeeOnRope
4y ago
Multiply is not as cheap as other arithmetic operations such as addition yet, though it has certainly gotten a lot cheaper (and many of these older bithack guides target CPUs that may not have a multiply at all). As an example, contemporary
108.
▲
by
BeeOnRope
4y ago
Many high performance locks do not, fundamentally, need to allocate memory. If they do, it is often for: 1) optional features such as statistics tracking or diagnostics (e.g., the ability to traverse all held locks in the process) 2) A spac
109.
▲
by
BeeOnRope
4y ago
> This has the whiff of someone discovering the basics of high-performance locking. Came here to say this. > So you need a startup calibration that measures it, otherwise it will have a random outcome. I guess one question is if AMD o
110.
▲
by
BeeOnRope
4y ago
You usually don't even need AVX512 to sustain enough load/stores at the core to max out memory bandwidth "in theory": even with 256 bit loads and assuming 2/1 loads/stores per cycle (ICL/Zen 3 and newer ca
111.
▲
by
BeeOnRope
4y ago
You were talking about reads and I responded about reads. I'll repeat my claim: SSDs largely behave as random access devices for page-aligned random reads of an integer multiple of pages. In particular, without any performance footguns
112.
▲
by
BeeOnRope
4y ago
Yes, Intel also takes a less than "full" approach to moving from 256b to 512. Though I think it is fair to say the Intel implementation represents kind of an intermediate state between the AMD approach (essentially no increase in
113.
▲
by
BeeOnRope
4y ago
SSDs read at page granularity. SSDs don't really have an equivalent to a HD seek: there is latency but this can over overlapped with the latency of dozens of other commands. Some SSDs may support reading an entire block in a faster way
114.
▲
by
BeeOnRope
4y ago
Yes, that's true though "blocks" is an overloaded term (e.g., people will talk about what "block size" they are reading at at the application layer). In any case I'm mostly disputing your characterization of SS
115.
▲
by
BeeOnRope
4y ago
That is not really true for SSD, which provide random access at the same performance, at least when a "suitable" block size is reached. This must be the case in fact because SSDs do not lay out data in the same order as the line
116.
▲
by
BeeOnRope
4y ago
SSDs generally read at the page level, so you can get full performance or close to it from much smaller reads than 256KB or more. For example, you may be able get maximum performance from 4K or 8K byte reads. Of course, there are a few la
117.
▲
by
BeeOnRope
4y ago
It's mentioned in the article (look for "What about the range-based for loop"): it works fine because it uses iterators under the covers which as local objects avoid the specific problem which occurs here. This article isn&#x
118.
▲
by
BeeOnRope
4y ago
Is that first example comparing squaring filing with the same _fixed_ rotation for each tile versus hex-tiling with random or otherwise varying rotation?
119.
▲
by
BeeOnRope
4y ago
Interestingly from the portability angle, this was all written in C# but with the SIMD intrinsics there he was still able to obtain ~SOTA performance. The repo I linked is his C++ version of the C# original. It was submitted to the C# stand
120.
▲
by
BeeOnRope
4y ago
No, because "freeing" memory makes sense within a single process and really means "removing the V->P mapping for the page", or more accurately something like "telling the malloc() implementation this pointer (imp
More ›