Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Const-me
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
61.
▲
by
Const-me
11mo ago
For performance critical code, you want to reuse L1D cache lines as much as possible. In many cases, allocation of a new immutable object boils down to malloc(). Newly allocated memory is unlikely to be found on L1D cache. OTOH, replacing d
62.
▲
by
Const-me
1y ago
> is written in languages that inherited their array ordering from C It’s not just C. Modern GPU hardware only supports row major memory layout for 2D and 3D textures (ignoring specialized layouts like swizzling and block compression but
63.
▲
by
Const-me
1y ago
What you wrote only applies to rotational latency, not seek latency. The seek latency is the time it takes for the head to reach the target. Heads only rotate within the small range like [ 0 .. 25° ], they are designed for rapid movements i
64.
▲
by
Const-me
1y ago
> would you know if there is something similar that works on u128 instead of just u32/u64? Not as far as I’m aware, but I think your use case is handled by the u64 version rather well. Instead of u128, use array of two uint64 intege
65.
▲
by
Const-me
1y ago
The support for BMI1 instruction set extension is almost universal by now. The extension was introduced in AMD Jaguar and Intel Haswell, both launched in 2013 i.e. 12 years ago. Instead of doing stuff like (word >> bit_offset) & s
66.
▲
by
Const-me
1y ago
> full-platter seek time: ~8ms; half-platter seek time (avg): ~4ms Average distance between two points (first is current location, second is target location) when both are uniformly distributed in [ 0 .. +1 ] interval is not 0.5, it’s 1&
67.
▲
by
Const-me
1y ago
For FFTW the showstopper was GPL license. For IPP, 200 MB of binary dependencies, also I remember when Intel was caught testing for Intel CPUs specifically in their runtime libraries instead or CPUID feature bits, deliberately crippling per
68.
▲
by
Const-me
1y ago
I have recently needed a decently performing FFT. Instead of doing Cooley-Tukey, I have realized the bruteforce version essentially computes two vector×matrix products, so I have interleaved and reshaped the matrices for sequential full-vec
69.
▲
by
Const-me
1y ago
Good article, but it uses less than ideal formula for weights of the gaussian blur kernel. Gaussian function for coefficients is fine for large sigmas, but for small blur radius you better integrate properly. Luckily, C++ standard library h
70.
▲
by
Const-me
1y ago
C# would catch the bug at compile time, just like Rust. https://www.rocksolidknowledge.com/articles/locking-asyncawa...
71.
▲
by
Const-me
1y ago
> does not account for frequency scaling on laptops Are you sure about that? > time spent in syscalls (if you don’t want to count it) The time spent in syscalls was the main objective the OP was measuring. > cycle counter While tec
72.
▲
by
Const-me
1y ago
Not sure if that’s relevant, but when I do micro-benchmarks like that measuring time intervals way smaller than 1 second, I use __rdtsc() compiler intrinsic instead of standard library functions. On all modern processors, that instruction m
73.
▲
by
Const-me
1y ago
Good article, but I believe it lacks information what specifically these magical dFdx, dFdy, and fwidth = abs(dFdx) + abs(dFdy) functions are computing. The following stackexchange answer addresses that question rather well: https:/&#
74.
▲
by
Const-me
1y ago
Interestingly, a few months ago I wanted a similar thing for Windows. Ended up developing a simple tray utility for that. Probably the most important method is the handler of WM_POWERBROADCAST message: https://github.com/Con
75.
▲
by
Const-me
1y ago
For the last few years, I’ve been developing CAM/CAE software on my job, sometimes embedded Linux. Same experience: last time I developed software written entirely in C++ was in 2008. Nowadays, only using C++ for DLLs consumed by C#, f
76.
▲
by
Const-me
1y ago
I believe Java is more popular for enterprise and web apps. .NET is widely used for videogames (games made with Unity, Godot, Unigine engines, also internal tools and game servers), desktop software, embedded software. Java is rarely used f
77.
▲
by
Const-me
1y ago
> Are there APIs which can sidestep the "load to CPU RAM" part? On windows that API is Desktop Duplication. The API delivers D3D11 textures, usually in BGRA8_UNORM format. When HDR is enabled you would need slightly different A
78.
▲
by
Const-me
1y ago
For the sending side, a small buffer (which should implement a flush method called after serializing the complete message) indeed helps amortize costs of system calls. However, a buffer large enough for 1GB messages will waste too much memo
79.
▲
by
Const-me
1y ago
> you never serialize a byte by byte over the network I sometimes do for the following two use cases. 1. When the protocol delivers real-time data, I serialize byte by byte over the network to minimize latency. 2. When the messages in qu
80.
▲
by
Const-me
1y ago
> Actually there is - you can exploit the data parallelism That doesn’t help much when you’re streaming the bytes, like many parsers or de-serializers do. You have to read bytes from the source stream one by one because each of the next
81.
▲
by
Const-me
1y ago
Great idea. BTW, when I implementing similar logic in C#, I place expressions like ((xored - 0x01010101U) & ~xored & 0x80808080U) inside unchecked blocks. Because when I compile my C#, I set <CheckForOverflowUnderflow>true<
82.
▲
by
Const-me
1y ago
> L1 cache is .. what.. 2 cycles? On Zen 4 CPU, I believe the typical latency of L1D is 4 cycles. However, if you (or your compiler) write AVX code which adds these floats, will still bottleneck on memory even if both inputs are in L1D c
83.
▲
by
Const-me
1y ago
LEB128 is slow and there’s no way around that. I don’t know why so many data formats are using that codec. The good variable integer encoding is the one used in MKV format: https://www.rfc-editor.org/rfc/rfc8794.html#na
84.
▲
by
Const-me
1y ago
I believe the reason why management specifically asked for animated GIFs was compatibility. The GIF format is ancient and supported by everything, makes it trivial to share or embed these animation files. Ignoring the compatibility, modern
85.
▲
by
Const-me
1y ago
I once did it for a closed-source CAM/CAE software. We wanted to generate high resolution and decent quality GIFs visualizing a progress of a numerical optimization in a few hundred animation frames. I wasn’t able to find a library whi
86.
▲
by
Const-me
1y ago
I’m not sure it’s register allocation. VC++ is indeed less than ideal, but 2x performance difference is IMO too much to explain by register allocation alone. First check that you’re passing correct flags to VC++ compiler and linker: optimiz
87.
▲
by
Const-me
1y ago
Indeed, automatic vectorizers do such simple things pretty reliably these days. However, if you build your software from kernels like that, you leave a lot of performance on the table. For example, each core of my Zen 4 CPU at base frequenc
88.
▲
by
Const-me
1y ago
The last few times I reserved a table at a restaurant, I only gave my first name. They never asked for my full name, would be hard for them to find me on social media. However, I'm an immigrant with a locally uncommon first name, YMMV.
89.
▲
by
Const-me
1y ago
> other systems had been exceeding 64 cores since the late 90s. Windows didn’t run on these other systems, why would Microsoft care about them? > x86 arguably didn't ship >64 hardware thread systems until then because NT didn&
90.
▲
by
Const-me
1y ago
The NT kernel dates back to 1993. Computers didn’t exceed 64 logical processors per system until around 2014. And doing it back then required a ridiculously expensive server with 8 Intel CPUs. The technical decision Microsoft made initially
More ›