Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mmozeiko
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Bringing Correctly Rounded Math to Production with LLVM-Libc
(devblogs.microsoft.com)
1 points
by
mmozeiko
16d ago
|
0 comments
2.
▲
Shared Memory Consistency from Scratch Part 1: Causality
(allthoughts.me)
2 points
by
mmozeiko
28d ago
|
0 comments
3.
▲
by
mmozeiko
2mo ago
See the following pdf for example on how to do this with SSSE3 (pages 104-133) or even SSE2 (pages 151-173) https://deplinenoise.files.wordpress.com/2015/03/gdc2015_afr...
4.
▲
by
mmozeiko
5mo ago
sub is also recognized as zeroing idiom for register file. Intel documents these in "3.5.1.7 Clearing Registers and Dependency Breaking Idioms" from Optimization Reference Manual: https://www.intel.com/content/
5.
▲
by
mmozeiko
6mo ago
xor swap trick was useful in older simd (sse1/sse2) when based on some condition you want to swap values or not: tmp = (a ^ b) & mask a ^= tmp b ^= tmp If mask = 0xfff...fff then a/b will be swapped, otherwise if ma
6.
▲
by
mmozeiko
7mo ago
How does it compare to rendering SVG by Direct2D itself? When using ID2D1DeviceContext5::DrawSvgDocument method, and ID2D1SvgDocument can be loaded from file with ID2D1DeviceContext5::CreateSvgDocument + SHCreateStreamOnFileW.
7.
▲
An Innocuous Blog Post about vPMU in QEMU
(vulpinecitrus.info)
3 points
by
mmozeiko
8mo ago
|
0 comments
8.
▲
by
mmozeiko
1y ago
Ah, I see. Yeah, then the current approach is fine.
9.
▲
by
mmozeiko
1y ago
From what I understand from code those unwarps are just doing matrix multiply to get unwraped pixel location? In this case doing these operations directly in fragment shader instead of texture lookup will be faster. Memory bandwidth is not
10.
▲
by
mmozeiko
1y ago
There's also Windows.Graphics.Capture. It allows to get texture not only for whole desktop, but just individual windows.
11.
▲
by
mmozeiko
1y ago
I don't know why Signal calls it "DRM" because the do not use DRM for this. Typically DRM means encryption & keys are involved (which is what Netflix & others are doing with Widevine or PlayReady). All Signal does is
12.
▲
by
mmozeiko
1y ago
If you change logic and/or to bitwise and/or then it'll be branchless.
13.
▲
by
mmozeiko
2y ago
https://github.com/veluca93/fpnge is a very fast png encoder. A bit lower compression ratio, but runs significantly faster than alternatives. Here is a presentation with benchmarks: https://www.lucaversari.i
14.
▲
by
mmozeiko
2y ago
There is a simple way to get that immediate from expression you want to calculate. For example, if you want to calculate following expression: (NOT A) OR ((NOT B) XOR (C AND A)) then you simply write ~_MM_TERNLOG_A | (~_MM_TE
15.
▲
by
mmozeiko
2y ago
Unfortunately not everything. For example, no xbox 360 controller support: https://github.com/microsoft/GDK/issues/39
16.
▲
by
mmozeiko
3y ago
It is called far2l: https://packages.debian.org/search?keywords=far2l&searchon=n...
17.
▲
by
mmozeiko
3y ago
Alt + start typing jumps to file/folder with this name. Ctrl + enter puts currently selected file/folder name into command-line. Ctrl + O hides panels and shows you "background" with all the previous command outputs. Cus
18.
▲
by
mmozeiko
3y ago
I recommend Polygon as virtual filesystem for SQLite database files: https://plugring.farmanager.com/plugin.php?l=en&pid=973 PortaDev for accessing MTP mounts (like Android): https://plugring.farmanager.com&#
19.
▲
by
mmozeiko
3y ago
On Ubuntu this seems to be official (?) ppa: https://launchpad.net/~far2l-team/+archive/ubuntu/ppa Make sure "far2l-gui" is installed. On ArchLinux there's package in AUR: https://au
20.
▲
by
mmozeiko
3y ago
Make sure you run wxgtk build, not the terminal one. Terminal is limited on what shortcuts it can do. But wxgtk build runs exactly as windows counterpart - Alt+F1/F2 and many other shortcuts work fine.
21.
▲
by
mmozeiko
3y ago
Here's a fancy trick from LLVM source: https://github.com/llvm/llvm-project/blob/main/llvm/lib/Targ... #define A 0xf0 #define B 0xcc #define C 0xaa And then you can build immedi
22.
▲
by
mmozeiko
3y ago
There's also dxwrapper that implements old DirectDraw/3D interfaces with newer D3D9 api for better compatibility with newer Windows versions: https://github.com/elishacloud/dxwrapper
23.
▲
by
mmozeiko
3y ago
I'm pretty sure such thing is implemented by hooking OS functions that are dealing with directory enumeration, create/open file, etc - and then redirecting to custom code that provides virtual files & their contents. There
24.
▲
by
mmozeiko
4y ago
Not fully open-sourced. There are bunch of pieces that are available only in compiled form - .obj/lib files. Like math.h functions.
25.
▲
by
mmozeiko
4y ago
You can use it for any workload that can use D3D12 buffers or textures as input. All the API does for you is transfer data from disk to ID3D12Resource object. After that it is up to you to do whatever you want - use it for fragment shader o
26.
▲
by
mmozeiko
4y ago
It is there to process last <32 elements. The vectorized loop processes up to 32 elements per iteration. The iteration does not happen if there are less than 32 elements left, because it wants to load 32 bytes as input. This is very typi
27.
▲
by
mmozeiko
4y ago
Btw there are alternative flac encoders, like FLACCL using GPU: http://cue.tools/wiki/FLACCL It compresses much faster than software libflac and gives smaller output files.
28.
▲
by
mmozeiko
5y ago
That's not really what's happening with such C++ code. `static` variables are placed in normal writeable global memory. Only difference with such static is that compiler generates extra "bool" variable and checks it to i
29.
▲
FJXL and FPNGE – Fast SIMD lossless image encoders [pdf]
(lucaversari.it)
1 points
by
mmozeiko
5y ago
|
0 comments
30.
▲
by
mmozeiko
5y ago
To me it looks like something related to some other optimization pass (I don't know much about gcc passes). But not related to writes to memory. Here are two writes both using cmov (on different code): https://godbolt.org&#x
More ›