Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
celrod
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
celrod
2y ago
Can confirm, this works for me in my actual examples, thanks!
32.
▲
by
celrod
2y ago
I've defined a few pretty printers, but `operator[]` doesn't work for my user-defined types. Knowing it works for vectors, I'll try and experiment to see if there's something that'll make it work. (gdb) p unroll
33.
▲
by
celrod
2y ago
How feasible would it be for something like gdb to be able to use a C++ interpreter (whether icpp, or even a souped up `constexpr` interpreter from the compiler) to help with "optimized out" functions? gdb also doesn't handle
34.
▲
by
celrod
2y ago
Yes, I see now that while not advertised on seller's websites, Asus's product pages do indeed say that.
35.
▲
by
celrod
2y ago
Skymont little cores have 4x 128-bit execution. They could quadruple-pump. But looks more like they're giving up on people writing code for wide vectors, instead settling on trying to make the existing code faster.
36.
▲
by
celrod
2y ago
Zen5 appears to officially support up to DDR5 5600, but unfortunately all of the ASRock Rack or Supermicro boards I looked at only supported DDR5 5200. I may wait for new Zen5 boards, or maybe take a gamble on something like the Asus ProArt
37.
▲
by
celrod
2y ago
Any suggestions for ECC? Would you suggest going with an ASRock Rack motherboard, even for desktop use, like you used here? https://www.phoronix.com/review/amd-ryzen9-ddr5-ecc I'm strongly tempted to get a Zen5 CP
38.
▲
by
celrod
2y ago
Signed integer overflow being undefined has these two consequences for me: 1. It makes my code slightly faster. 2. It makes my code slightly smaller. 3. It makes my code easier to check for correctness, and thus makes it easier to write cor
39.
▲
by
celrod
2y ago
I don't think you'd even necessarily need to ignore. Roll it out in phases. You aren't going to have to deliver the final finished solution all at once. Some elements are inevitably going to end up being de-prioritized, and p
40.
▲
by
celrod
2y ago
Thanks for the clarification.
41.
▲
by
celrod
2y ago
C++20 added `[[no_unique_address]]`, which lets a `std::is_empty` field alias another field, so long as there is only 1 field of that `is_empty` type. https://godbolt.org/z/soczz4c76 That is, example 0 shows 8 bytes, f
42.
▲
by
celrod
2y ago
Multiple accumulators increases accuracy. See pairwise summation, for example. SIMD sums are going to typically be much more accurate than a naive sum.
43.
▲
by
celrod
2y ago
C++23 added `allocate_at_least`: https://en.cppreference.com/w/cpp/memory/allocator_traits/al... I'm not sure if any standard libraries have an implementation that takes advantage of the "at le
44.
▲
by
celrod
2y ago
Yes! It annoys me when a scene with characters shouting is much louder than a scene where characters are talking with hushed voices, as an example. We know a shout was louder at the source, but the decibel level at our ears is proportional
45.
▲
by
celrod
2y ago
Nontemporal writes are substantially slower, e.g. with avx512 you can do 1 64 byte nontemporal write every 5 or so clock cycles. That puts you at >= 640 cycles for 8 KiB. https://uops.info/html-instr/VMOVNTPS_M512_ZM
46.
▲
by
celrod
2y ago
This memory is now the least recently used in the L1 cache, despite being freed by the allocator, meaning it probably isn't being used again. If it was freed after already being removed from the L1 cache, then you also need to evict ot
47.
▲
by
celrod
2y ago
Different vector widths for different cores isn't currently feasible, even with SVE. So all cores would need to support 1024-bit SIMD. I think it's reasonable for the non-SIMD focused cores to do so via splitting into multiple mic
48.
▲
by
celrod
2y ago
I read that comment as "the wider, the sweeter" (which I agree with), but that we're now (as you say) at the end of the road, and thus the sweetest point. But an increase in cacheline size would be nice if it can get us large
49.
▲
by
celrod
2y ago
Our software ecosystem doesn't work well with an army of ants. I think we'd need a paradigm shift to get there. Also, FWIW, Xeon Phi hit 244 threads in 2012 and 256 threads in 2016, although it used 4 threads/core.
50.
▲
by
celrod
2y ago
Yeah, I don't think it's useful (except to score [office] political points) to read the least generous interpretation you can. Trying to understand a position lets you better decide whether or not it actually makes sense, e.g. Che
51.
▲
by
celrod
2y ago
> We live in mortal fear of compiler writers smiting us for innocent things like punning through a union. C++20 introduced `std::bitcast`, so I appreciate alias analysis getting all the help it can.
52.
▲
by
celrod
2y ago
Yes, I have an AVX512 double precision exp implementation that does this thanks to iperm2pd. This approach was also recommended by the Intel optimization manual -- a great resource. I just went with straight math for single-precision, thoug
53.
▲
by
celrod
2y ago
The Apple M4 is rumored to be ARMv9, even featuring SME on top of SVE2: https://wccftech.com/apple-m4-adopts-armv9-run-complex-workl... If true, I'd much rather buy an M4 and use Asashi linux (once it supports the M4)
54.
▲
by
celrod
2y ago
Bill Dally from Nvidia argues that there is "no gain in building a specialized accelerator", in part because current overhead on top of the arithmetic is in the ballpark of 20% (16% of IMMA and 22% for HMMA units) https:/&#x
55.
▲
by
celrod
2y ago
EXWM is ;). I used it for about a year, but ultimately found the performance, and in particular a low language server locking up the window manager, unacceptable.
56.
▲
by
celrod
2y ago
I'm still waiting for clangd support, e.g. [0] before trying modules. But maybe I should just try it, as at least one person reports that it already works [1]. [0] https://github.com/clangd/clangd/issues/
57.
▲
by
celrod
2y ago
What's the trick with explicit template instantiations? Including them in the precompiled header?
58.
▲
by
celrod
2y ago
If using cmake, you can try a unity build. https://cmake.org/cmake/help/latest/prop_tgt/UNITY_BUILD.htm... You can also specify a `-DUNITY_BUILD_BATCH_SIZE` to control how many get grouped, so you can st
59.
▲
by
celrod
3y ago
FWIW, that synthetic benchmark was reflective of some real world code we were deploying. Using malloc/free for one function led to something like a 2x performance improvement of the whole program. I think it's important to differe
60.
▲
by
celrod
3y ago
The RCU use case is convincing, but my experience with GCs in other situations has been poor. To me, this reads more like an argument for bespoke memory management solutions being able to yield the best performance (I agree!), which is a to
More ›