Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Nyan
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
31.
▲
by
Nyan
6y ago
Thanks for the info! Unfortunately I don't have access to any VIA/Centaur CPUs, so couldn't test on those (though test results welcome if anyone is willing/able to!). But yeah, you have to check the CPU you're runni
32.
▲
by
Nyan
6y ago
> Are there real uses for this kind of thing, on modern architectures? For me, I came up with an algorithm for doing error correction coding, however, good performance can only be achieved by JIT'ing code. Trying to implement the al
33.
▲
by
Nyan
6y ago
Thanks for the info. I'm not particularly familiar with common JIT applications, but I suspect that this use-case is actually more niche than may think. The problem is that the example presented requires a memory page with write + exec
34.
▲
Single-use JIT Performance on x86 Processors
(github.com)
92 points
by
Nyan
6y ago
|
25 comments
35.
▲
by
Nyan
7y ago
> I would almost prefer a more predictable, high-latency decomposition into 4x128 wide uops over what we have now. AVX512-VL gives the programmer AVX512 functionality at 128/256-bit widths, if it is believed to be more beneficial th
36.
▲
Fast Galois Field Region Multiplication Techniques
(github.com)
3 points
by
Nyan
7y ago
|
0 comments
37.
▲
by
Nyan
7y ago
Note that SIMD on CPUs is somewhat different to GPU style "SIMD" (particularly for cases where you want fixed width SIMD vs arbitrary vector processing (ala SPMD, ISPC, CUDA etc)). JSON parsing, for example, doesn't scale in
38.
▲
by
Nyan
7y ago
> It's hard to program, but maybe we need to figure that out in order to get performance. There will always be problems which require latency over throughput. Although the idea has been tried - e.g. Xeon Phi.
39.
▲
by
Nyan
7y ago
You're confusing instruction-set architecture with micro-architecture. The x86-64 ISA defines 16 integer ("logical") registers, for example, however, Skylake has (don't quote me on the figure) 168 physical integer regi
40.
▲
by
Nyan
9y ago
yEnc is very much designed specifically for Usenet though. It gets its low overhead largely from using almost all characters which aren't special in NNTP (and also adapts allowed characters depending on NNTP context), which isn't
41.
▲
by
Nyan
10y ago
Just FYI, here's an algorithm [ https://github.com/animetosho/ParPar/blob/master/xor_depends... ] which explicitly relies on JIT being a fast operation. It's different from your typical language