Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Const-me
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
15 ms
·
151.
▲
by
Const-me
2y ago
Yeah, every time I see articles about importance of linear color space for gradients, and see images there, I observe the opposite of what’s written in the text of these articles. Gradients in sRGB color space look better. I have a suspicio
152.
▲
by
Const-me
2y ago
Stay calm, and support armed forces of Ukraine if you can. With sufficient external support, Ukrainians will make Russia small again. Personally I trust the following two foundations: https://savelife.in.ua/en/ https:
153.
▲
by
Const-me
2y ago
> Obviously, you can't use it for embedded Embedded is diverse. I would not use .NET for small embedded, i.e. stuff running on Arduino or ESP32. However, I have successfully used .NET runtime in production for embedded software runn
154.
▲
by
Const-me
2y ago
> the dividing line seems to be at 1e33.. Not sure what to make of that That’s not too bad. They are probably using hand-rolled FP128 format for their numbers. If they were using hardware-provided FP64 arithmetic, the threshold would hav
155.
▲
by
Const-me
2y ago
Due to backwards compatibility modern PC CPUs have some mathematical constants in hardware, one of them Pi https://www.felixcloutier.com/x86/fld1:fldl2t:fldl2e:fldpi:f... Moreover, that FLDPI instruction delivers 80 bi
156.
▲
by
Const-me
2y ago
> Once consolidation shrinks the number of players down to a handful or less, that's when competition and innovation are stifled. Yeah, and ads is the main mechanism why. Ban ads of carbonated drinks, and pepsi/coca cola will b
157.
▲
by
Const-me
2y ago
> competition drives improvement Are you saying ads helping competition? How? > customer choice is central to a functioning market, leading to innovation. Ads suppress customer choice, along with innovations. Can you recall any recent
158.
▲
by
Const-me
2y ago
> not sure how best to address this sort of thing I use two browser addons, uBlock origin and SponsorBlock.
159.
▲
by
Const-me
2y ago
The only way is do what OP did – compile your shader, disassemble, and read the assembly. I do that quite often with my HLSL shaders, learned a lot about that virtual instruction set. For example, it’s interesting GPUs have instruction sinc
160.
▲
by
Const-me
2y ago
> how does that hardware implementation work internally? I don’t know, but based on performance difference between FP32 and FP64 square root instructions, the implementation probably produces 4-5 bits of mantissa per cycle.
161.
▲
by
Const-me
2y ago
> at least for sqrt(), internally it's likely implemented as a heuristic guess Square roots are implemented in hardware: https://www.felixcloutier.com/x86/sqrtsd > In software, you'd normally iterate N
162.
▲
by
Const-me
2y ago
> Have you never downscaled and upscaled images in a non-3D-rendering context? Indeed, and I found that leveraging hardware texture samplers is the best approach even for command-like tools which don’t render anything. Simple CPU running
163.
▲
by
Const-me
2y ago
> Trilinear interpolates across three dimensions, such as 3D textures or mip chains I meant trilinear interpolation across the mip chain. > generated with "some" downsampling filter (which can be anything from box to Lanczos
164.
▲
by
Const-me
2y ago
Good article, but IMO none of the discussed downsampling methods are actually ideal. The ideal is trilinear sampling instead of bilinear. Easy to do with GPU APIs, in D3D11 it’s ID3D11DeviceContext.GenerateMips API call to generate full se
165.
▲
by
Const-me
2y ago
I believe it’s ecosystem. For example, Python is very high level language, not sure it even has a memory model. But it has libraries like NumPy which support all these vectorized exponents when processing long arrays of numbers.
166.
▲
by
Const-me
2y ago
Yeah, but if you need C versions of low-level SIMD algorithms from there, copy-paste the implementation and it will probably compile as C code. That’s why I mentioned the MIT license: copy-pasting from GPL libraries may or may not work depe
167.
▲
by
Const-me
2y ago
I would rather conclude that automatic vectorizers are still less than ideal, despite SIMD instructions have been widely available in commodity processors for 25 years now. The language is actually great with SIMD, you just have to do it yo
168.
▲
by
Const-me
2y ago
> an example of a communist state which came into existence and was not immediately beset by foreign intelligence Post-war Yugoslavia was such an example. As for the brutality, here’s a link: https://en.wikipedia.org/wiki
169.
▲
by
Const-me
2y ago
I’m not sure it matters; I hope market economy should do the rest. I believe it’s slightly cheaper to manufacture a device with a single USB-C port which supports PD, compared to a device with two ports, one 5W USB-C and some other port for
170.
▲
by
Const-me
2y ago
> I've tried it three times Have you tried compute shaders instead of that weird HPC-only stuff? Compute shaders are widely used by millions of gamers every day. GPU vendors have huge incentive to make them reliable and efficient: m
171.
▲
by
Const-me
2y ago
PyTorch is a project by Linux foundation. The about page with the mission of the foundation contains phrases like “empowering generations of open source innovators”, “democratize code”, and “removing barriers to adoption”. I would argue run
172.
▲
by
Const-me
2y ago
Unlike training, ML inference is almost always bound by memory bandwidth as opposed to computations. For this reason, tensor cores, cuDNN, and other advanced shenanigans make very little sense for the use case. OTOH, general-purpose compute
173.
▲
by
Const-me
2y ago
It’s not terribly hard to port ML inference to alternative GPU APIs. I did it for D3D11 and the performance is pretty good too: https://github.com/Const-me/Cgml The only catch is, for some reason developers of ML libra
174.
▲
by
Const-me
2y ago
> What do you mean by "something completely different"? I mean compute the exact thing you need to compute, optimizing for performance. I’m pretty sure your current code is spending vast majority of time doing malloc / mem
175.
▲
by
Const-me
2y ago
> until you get into the millions of records per second level, you're almost never benefited Yeah, but the software landscape is very diverse. On my job (CAM/CAE) I often handle data structures with gigabytes of data. Worse, un
176.
▲
by
Const-me
2y ago
The input data is sequential. I don’t understand Rust but if the code is doing what’s written there, “simple multiplicative hash and perform a simple analysis on the buckets – say, compute the sum of minimums among buckets” it’s possible to
177.
▲
by
Const-me
2y ago
In my 3 years old laptop, system memory (dual channel DDR4-3200) delivers about 50 GB / second. You have measured 50M elements per second, with 8 bytes per element translates to 400 MB / second. If your hardware is similar, your i
178.
▲
by
Const-me
2y ago
> people like you are a tiny minority According to some rumors, Valve is currently developing a Linux-based computer with software stack ported from Steam Deck. If the rumor is true and the device will be good, Microsoft will be very sur
179.
▲
by
Const-me
2y ago
> CPU-only custom 2D pixel blitter engine I wrote to make 2D games in styles practically impossible with modern GPU-based texture rendering engines I’m curious what’s so special about that blitting? BTW, pixel shaders in D3D11 can receiv
180.
▲
by
Const-me
2y ago
I wonder how does the perf in tokens/second compares to my version of Mistral: https://github.com/Const-me/Cgml/tree/master/Mistral/Mistral... BTW, see that section of the readme about quantiza
More ›