Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
exDM69
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
61.
▲
by
exDM69
1y ago
> How exactly does your debugger know whether the compiled code it stepped into came from C++ or Fortran source? Executables with debug symbols contain the names of the source files it was built from. Your debugger understands the debug
62.
▲
by
exDM69
1y ago
GPU texture compression needs fixed bitrate for random access, usually 2 bits per pixel (128b for 8x8 block) with modern formats. Jpeg or other variable bitrate compression schemes are not applicable.
63.
▲
by
exDM69
1y ago
It's an open ended optimization problem. You can get bad results very quickly but on high quality settings it can take minutes or hours to compress large textures. And the newer compression formats have larger block sizes from 8x8 to e
64.
▲
by
exDM69
1y ago
Note that KTX2 is a container format. It can contain uncompressed or compressed texture data. Some of the tools you mention might support "KTX2" but they won't compress textures for you, they just put the uncompressed data in
65.
▲
by
exDM69
1y ago
Not for multi-channel SDF at least. Texture compression works terribly badly with "uncorrelated" RGB values as they work in chroma/luminance rather than RGB. For uncorrelated values like normal maps, there are texture compres
66.
▲
by
exDM69
1y ago
Has anyone got suggestions for blur algorithms suitable for compute shaders? The usual Kawase blur (described in the article) uses bilinear sampling of textures. You can, of course, implement the algorithm as is on a compute shader with tex
67.
▲
by
exDM69
1y ago
The article doesn't describe the way to avoid the difference in rectangle size in a cubesphere, so let me. The bad way: - Generate a cube - Subdivide each face using linear interpolation (lerp) - Normalize each vector to put it on a un
68.
▲
by
exDM69
1y ago
The system package manager and the language package/dependency managers do a very different task. The distro package manager delivers applications (like Firefox) and a coherent set of libraries needed to run those applications. Most di
69.
▲
by
exDM69
1y ago
The same applies to any Makefile, the Python script invoked by CMake or pretty much any other scriptable build system. They are all untrusted scripts you download from the internet and run on your computer. Rust build.rs is not really spec
70.
▲
by
exDM69
1y ago
It's not always necessary to check for inf/NaN explicitly using isinf/isnan. Both inf and NaN are floating point values with well defined semantics. I'll give two examples from a recent project where I very intentionally
71.
▲
by
exDM69
1y ago
My favorite thing about floating point numbers: you can divide by zero. The result of x/0.0 is +/- inf (or NaN if x is zero). There's a helpful table in "weird floats" [0] that covers all the cases for division an
72.
▲
by
exDM69
1y ago
> waiting required is of the form "X can't proceed until A, B, D, and Q are done In my experience, this is the common pattern on GPU workloads. On the CPU (where async happens), the wait pattern is usually much simpler. Of cour
73.
▲
by
exDM69
1y ago
> Also, from a strictly prose point of view, isn't it strange that the `clz` instruction It's using the `bsr` instruction which is similar (but worse). The `lzcnt` instruction in x86_64 is a part of the BMI feature introduced i
74.
▲
by
exDM69
1y ago
> Lots of Vulkan map quite well to Rust's ownership rules, the memory allocation API surface maps very well. But anything that's happening on the GPU timeline is pretty much impossible to do safely. I agree with this, having be
75.
▲
by
exDM69
1y ago
The article that GP posted was specifically about throughput over a high speed connection inside a data center. It was not about latency. In my opinion, the lessons that one can draw from this article should not be applied for use cases tha
76.
▲
by
exDM69
1y ago
No, it does not. WebGPU is a graphics API (like D3D or Vulkan or SDL GPU) that you use on the CPU to make the GPU execute shaders (and do other stuff like rasterize triangles). Rust-GPU is a language (similar to HLSL, GLSL, WGSL etc) you ca
77.
▲
by
exDM69
1y ago
Yes it can and yes indeed it's sometimes a pain in the ass. However, some (maybe even most) C libraries' Rust bindings deal with this automatically when installed via Cargo.
78.
▲
by
exDM69
1y ago
Long distance communication with submarines is difficult but not impossible, there are four extra low frequency transmitters in the world to send signals in nuclear doomsday scenarios. For shorter distances, there are acoustic and optical d
79.
▲
by
exDM69
1y ago
It's actually Rust that is setting the 64 wide limit (see SupportedLaneCount), not LLVM. I agree, f64x64 is probably a very bad idea. But something like f32x8 would probably still be "fast enough" on old/mobile CPUs with
80.
▲
by
exDM69
1y ago
Your experience matches mine, you can get a lot done with the portable SIMD in Clang/GCC/Rust but you can't avoid the platform specific stuff when you need specialized instructions. Depends on the domain you work in how much
81.
▲
by
exDM69
1y ago
> Practically, it really is hard, because SIMD instruction sets in CPUs are a mess. X86 and ARM have completely different sets of things that they have instructions for Not disagreeing it's a mess, but there's also quite a big
82.
▲
by
exDM69
1y ago
> if you don't describe your code and dataflow in a way that caters to the shape of the SIMD But when I do describe code, dataflow and memory layout in a SIMD friendly way it's pretty much the same for x86_64 and ARM. Then I ca
83.
▲
by
exDM69
1y ago
Here's a toy example of f32 vs. f64 generic Rust std::simd code without macros. It aint't pretty but it works. https://play.rust-lang.org/?version=nightly&mode=debug&editi... In my projects, I've put
84.
▲
by
exDM69
1y ago
> f64x64 is presumably an exaggeration It's not. IIRC, 64 elements wide vectors are the widest that LLVM (or Rust, not sure) can work with. It will happily compile code that uses wider vectors than the target CPU has and split accor
85.
▲
by
exDM69
1y ago
Here we are discussing the merits of built in SIMD facilities of not one but two programming languages. Waiting for the Zig guys to chime in to make it a three. Four if you include LLVM IR (I don't). No reason to be dismissive about it
86.
▲
by
exDM69
1y ago
The binary portability issue is not specific to Rust or std::simd. You would have to solve the same problems even if you use intrinsics or C++. If you use 512 but vectors in the code, you will need to check if the CPU supports it or add mul
87.
▲
by
exDM69
1y ago
This right here illustrates why I think there should be better first class SIMD in languages and why intrinsics are limited. When using GCC/clang SIMD extensions in C (or Rust nightly), the implementation of sin4f and sin8f are line by
88.
▲
by
exDM69
1y ago
Admirable perseverance! I've always also had a side project or two in this domain but I've never managed to stick with one for more than 3-5 years.
89.
▲
by
exDM69
1y ago
> How would you use shared/local memory in GLSL? In compute shaders the `shared` keyword is for this.
90.
▲
by
exDM69
1y ago
Thanks for the words of encouragement. I need more time to turn this prototype into a practical piece of software and/or a publication (blog post or research paper). Unfortunately there are only so many hours in a day. To actually rend
More ›