Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
camel-cdr
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
camel-cdr
2mo ago
> And maybe 16-32-48-64 is worth it.. It is, with a prefix encoding, you can reuse the RVC decode path 1-to-1 and get the 48/64-bit instruction starts with a simple bitshift (or simply handle the 48/64-bit instructions via the
32.
▲
by
camel-cdr
2mo ago
> The fact that it's only "competitive" with aarch64's code density is a solid black mark against RISC-V. Arm uses complex instructions with multiple writeback, that require cracking, to improve code density. RISC-V u
33.
▲
by
camel-cdr
2mo ago
Nobody in high-performance does fixed-width instructions that allow lineary scaling parallel decoders. Arm basically requires certain instructions to be cracked into multiple uops before rename. That ends up analougus to decoding compressed
34.
▲
by
camel-cdr
2mo ago
My disagreement with the article is mostly the following: RISC-V is not an ISA, but an ISA generation framework. If RISC-V would've standardized aarch64 1-to-1, the end result would've still been a huge extension mess, because a l
35.
▲
by
camel-cdr
2mo ago
My disagreement with the article is mostly the following: RISC-V is not an ISA, but an ISA generation framework. If RISC-V would've standardized aarch64 1-to-1, the end result would've still been a huge extension mess, because a l
36.
▲
by
camel-cdr
2mo ago
RISC-V has higher code density than Arm and x86 and all three have a comparible uop density, but RISC-V has a slightly higher average instruction count. What more complex instructions are missing?
37.
▲
by
camel-cdr
2mo ago
can you share the data? btw, how much protein do you consider recommended?
38.
▲
by
camel-cdr
2mo ago
I wonder what substitute the average vegan consumes. I often see these vegan meat substitute products as an on-ramp and people who are vegan for longer tend to use other things in their staple dishes.
39.
▲
by
camel-cdr
2mo ago
> twice the native vector width has ~20% better throughput Yes, this is what I was saying, but twice the vector width of AVX-512 will perform horrible in SSE, which is why portable SIMD abstractions should make writing code relative to t
40.
▲
by
camel-cdr
2mo ago
Except you can't use this in actual code, because either, as is the case in this example with f32x32, you run out of registers and spill all over the place. Or you aren't using your full vector register or could've gotten bet
41.
▲
by
camel-cdr
2mo ago
The create is called portable_simd. There is no reason a portable_simd relu_dot implemention should need to specify the SIMD width. But the design and documentation of portable_simd makes the fixed size syntactically easy/the default a
42.
▲
by
camel-cdr
2mo ago
I love how ever example of portable SIMD isn't portable. They specifies a constant SIMD width so it's non-portable. Well, not performance portable, but why are we using SIMD again?
43.
▲
Google search filter by date seems broken
5 points
by
camel-cdr
2mo ago
|
2 comments
44.
▲
by
camel-cdr
2mo ago
this type of thing usually means you are the product
45.
▲
by
camel-cdr
3mo ago
Well, this explains who has the training data to fake it then.
46.
▲
by
camel-cdr
3mo ago
It used to be $300 for the 8GB RAM version for the first batch in May/June. (16GB for $400)
47.
▲
by
camel-cdr
3mo ago
> So, why not just cover that flaw by adding a carry/overflow flag. Actually, why not just have a single generic 1-bit flag? You could still have cmp + branch, but now you put the comparison type into the cmp. This gives you a large
48.
▲
by
camel-cdr
3mo ago
> Well, cmov is absolutely possible without flags. You see such "select" instructions all the time on GPUs Well, yes, because GPUs need 3r1w anyways, for fmadd, and because a dedicated cmov (blen/merge) is a lot more impor
49.
▲
by
camel-cdr
3mo ago
Designing an ISA arround carry flags certainly gives you many goodies: cmov, adc, ccmp, larger branch range
50.
▲
by
camel-cdr
3mo ago
That seems like a weird thing to save on, for architectures that have expensive but rarely used things like CLZ on all ALUs. But I hadn't considered the ISA design option of only carry flag, no seperate cmp and branch, before.
51.
▲
by
camel-cdr
3mo ago
> Out-of-order pipelines actually have a really elegant way of handling flags, they just store a copy of the flags register on every ROB entry. And yet, flag writing instructions are usually half the throughput on regular ALU instruction
52.
▲
by
camel-cdr
3mo ago
No don't buy this one, it doesn't support RVV. Their newer ones do, but those are a lot more expensive.
53.
▲
by
camel-cdr
3mo ago
> Dart supports RISC-V as a target architecture for compilation, but I'm not really excited about figuring out how to map the wasm-SIMD-style primitives to RISC-V's RVV and so I don't really plan to look into it at all. On
54.
▲
by
camel-cdr
3mo ago
> No, not really. You can still jump into the middle a 32-bit instruction, and it's possible it can be reinterpreted as a valid 32/16-bit instruction. Remember when people complained about how "overlapped instructions"
55.
▲
by
camel-cdr
3mo ago
You can see the encoding limitation it in the design on SVE, which only has destructive operations, but MOVPRFX, which is a round about way of doing 64-bit instructions, without doing 64-bit instructions.
56.
▲
by
camel-cdr
3mo ago
You can benchmark stuff with and without RVC. Once there is faster hardware I want to do such a comparison with a full gentoo build for both sides. However, quantifying what the result will actually mean is nearly impossible, because you do
57.
▲
by
camel-cdr
3mo ago
> E.g. you fetch 66 bytes instead of 64 Not really, you would fetch fewer bytes with RVC [2, page 9], because the code density is better. > I would be really surprised if the lower code density is worse than the improvement due to eve
58.
▲
by
camel-cdr
3mo ago
There are two options when designing an ISA to achieve competitive code size, add variable length instructions or add more complex fixed-length instructions which require cracking (2W instructions). The other option is: maybe codesize don&#
59.
▲
by
camel-cdr
3mo ago
> Yes that's precisely the point of excluding it from the RVA profiles. It would mean that Linux distros don't compile code with C enabled, so chips are free to not support C and therefore can achieve higher performance (probab
60.
▲
by
camel-cdr
4mo ago
\* Fibonacci hashing spreads packed ARGB keys uniformly. Used so that low bits don't dominate the cache. */ return (NSUInteger) ((key * 11400714819323198485ULL) & (NS_COLOR_CACHE_SIZE - 1)); Th
More ›