Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mshockwave
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
mshockwave
1y ago
This is pretty useful! Any plan for adding ARM SVE and RISC-V V extension?
32.
▲
by
mshockwave
1y ago
> that modeling instruction scheduling doesn't matter all that much for codegen on OoO cores. yeah scheduling quality usually has a weaker connection to the performance of OoO cores. Though I would also like to point out: 1. in-
33.
▲
by
mshockwave
1y ago
> This means I can't feed any function as-is (even a tiny one). I need to manually cut out the loop body. > It doesn't support branches at all. I know it's a very hard problem, but that's the problem I have Shamele
34.
▲
by
mshockwave
1y ago
It was later reverted[1] because "there are devices in the field using usbX interfaces for tethering". Shortly after that, it got re-landed but only supported Android V+[2] [1]: https://android-review.googlesource.com&#
35.
▲
by
mshockwave
1y ago
Nitpick: static link won’t give you inlining but only eliminates the overhead of PLT. LTO will give you more opportunities for inlining
36.
▲
by
mshockwave
1y ago
> In the case where one of the structs has just been computed, attempting to load it as a single 32-bit load can result in a store forwarding failure It actually depends on the uArch, Apple silicon doesn't seem to have this restrict
37.
▲
by
mshockwave
1y ago
sometimes also known as superoptimization, which many of them also use SMT solvers like Z3 mentioned in the article
38.
▲
by
mshockwave
1y ago
I might be wrong, but it seems like Anubis asks the _client_ to solve the cryptography challenges while the approach Cloudflare described here asks the server to verify the (cryptography) signature?
39.
▲
by
mshockwave
1y ago
in addition to storing profiles, what about caching some native code? so that we can eliminate the JIT overhead for hot functions EDIT: they describe this in their "Alternative" section as future work
40.
▲
by
mshockwave
1y ago
judging from some of these repos, I think the answer is: they don't. It seems like you have to manually pick images with the correct resolution
41.
▲
by
mshockwave
1y ago
nit: it shouldn’t be RV64GC if it has vector extension
42.
▲
Land ahoy: leaving the Sea of Nodes
(v8.dev)
3 points
by
mshockwave
1y ago
|
0 comments
43.
▲
by
mshockwave
2y ago
I think this is exactly how Zig compiler does under the hood for C/C++ sources. So I guess you can do a similar thing for your own programming languages to support interop.
44.
▲
by
mshockwave
2y ago
yes, almost all DSPs I know have native HW supports for FFT, since it's the bread and butter for signal processing
45.
▲
by
mshockwave
2y ago
It’s pretty common to roll your own linker script in embedded software development
46.
▲
by
mshockwave
2y ago
> Here's my question: It seems inevitable that people will eventually port all LLVM's smarts directly into MLIR, and remove the need to shift between the two. Is that right? Theoretically, yes -- taking years if not decades for
47.
▲
by
mshockwave
2y ago
modular in terms of using only some of the LLVM libraries without the need to pull the entire compiler into your project. In fact, many of the LLVM libraries have absolutely nothing to do with LLVM IR and have zero dependency on it. For ins
48.
▲
by
mshockwave
2y ago
I’m definitely happy to see this happening. But I would like to point out two ingredients that constitute LLVM’s success beyond academic merits: License and modularity. I’m not a lawyer so can’t say much about the first one, all I can say i
49.
▲
Visualize RISC-V Vector Memory Instructions
(myhsu.xyz)
4 points
by
mshockwave
2y ago
|
0 comments
50.
▲
by
mshockwave
2y ago
> ALUs in Recent Apple cpus can actually start a new division every other cycle (in addition to having an abnormally low latency) That's indeed impressive. I'll argue that we're definitely capable of making fully pipelined
51.
▲
by
mshockwave
2y ago
shameless self plug of modern uArchs and how LLVM models it: https://myhsu.xyz/llvm-sched-model-1
52.
▲
by
mshockwave
2y ago
divisions, regardless of integer or floating point, are usually NOT pipelined though
53.
▲
by
mshockwave
2y ago
> Of course I have to test it with https://gitdiagram.com/torvalds/linux I really wish to see how well (or bad) it works on mega projects. Because those are usually the ones I need diagrams like this the most.
54.
▲
by
mshockwave
2y ago
Zb extension is in both RVA22 and RVA23 profiles, meaning application cores (targeting consumer devices like smartphones) designed in the past few years almost certainly have shXadd instructions in order to be compatible with the mainstream
55.
▲
by
mshockwave
2y ago
SpacemiT K1 on BananaPi is another commonly seen RVV 1.0 capable chip. IIRC both Kendryte K230 and SpacemiT K1 are in-order cores.
56.
▲
by
mshockwave
2y ago
> 2 others that I honestly have no idea what they do CSR (i.e. status) register and instruction fence extensions. Instruction fences are most useful in cases where you modify text section during runtime (e.g. JIT or code hot reload) such
57.
▲
by
mshockwave
2y ago
Re: Why does link-time optimization (LTO) happen at link-time? I think maybe LLVM's ThinLTO is what you're looking for where whole-program optimization (more or less) happens in middle end.
58.
▲
by
mshockwave
2y ago
> As an example, in my desktop, it takes ~25min to compile a Debug version of LLVM, but ~23min to compile a Release version. oh I think I know what might cause this: TableGen. The `llvm-tblgen` run time accounts for a good chunk of LLVM
59.
▲
by
mshockwave
2y ago
> The compiler optimizes for data locality > So, we have a single array in which every entry has the key and the value paired together. But, during lookups, we only care about the keys. The way the data is structured, we keep loading
60.
▲
by
mshockwave
2y ago
I wouldn't recommend using the term "devirtualization" here, as that term has been used to refer simplifying C++ virtual function calls (into normal function call) in LLVM. And such optimization has been turned on by default
More ›