Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Asm2D
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
Asm2D
3y ago
Can you elaborate what's "bad" on asmjit? I wrote it for the purpose of making it easy to write JIT compilers in C++ and it has been adopted even by projects that are written in C (like Erlang). It's also one of the smal
32.
▲
by
Asm2D
3y ago
They do! The most important post-SSE2 extensions are SSSE3 (pshufb) and SSE4.1 (rounding, min/max, blending, etc...). Pure SSE2 is a nightmare to use as it's a totally unbalanced SIMD ISA (a lot of missing stuff here are there req
33.
▲
by
Asm2D
3y ago
It's above 4400 instructions actually if you count different encoding of SIMD instructions (like SSE2, VEX, and EVEX variations) and consider instructions using 128-bit, 256-bit, and 512-bit SIMD as separate instructions. Instructions
34.
▲
by
Asm2D
3y ago
A C version of the library with a portable implementation and AVX-512 optimizations is planned.
35.
▲
by
Asm2D
3y ago
The initial AVX-512 implementation brought a lot of issues with it. The biggest problem was that Intel used 512-bit ALUs from the beginning and I think it was just too much that time (initial 14nm node) - even AMD's Zen4 architecture,
36.
▲
by
Asm2D
4y ago
Competing solutions are great for innovation.
37.
▲
by
Asm2D
4y ago
This is good for entertainment purposes as it shows how to encode individual instructions and how to play with mmap() to actually execute the generated code. However, for anything practical I would recommend AsmJit as it offers a lot of fea
38.
▲
by
Asm2D
5y ago
This reminds me my older project called xql, for node.js: https://github.com/jsstuff/xql and fiddle: https://kobalicek.com/experiments/fiddle-xql.html Interesting how all these builders look the s
39.
▲
by
Asm2D
6y ago
Ok, so you cannot prove it, which means that I must prove otherwise? I wrote exactly what you cited in May 2019, but it was not about SKIA. And if you continued reading you would have noticed that multi-threaded rendering context implementa
40.
▲
by
Asm2D
6y ago
Do you think it's right to state something because of "common wisdom"? Claims like "order of magnitudes faster" should be supported by a reference, because from my own experience making something order of magnitude
41.
▲
by
Asm2D
6y ago
Can you post a link to any benchmark that would prove that?
42.
▲
by
Asm2D
6y ago
There is an aarch64 branch in AsmJit project that provides an experimental AArch64 backend. It's pretty complete. Having a platform independent IR in AsmJit is not planned. AsmJit is more about control and ISA completeness. IR could of
43.
▲
by
Asm2D
6y ago
AsmJit was designed to be able to integrate well with C code bases. It uses C-style error handling (no exceptions) and provides easy to use API. I don't think it's a big deal to use it in a C project - there are other C projects t
44.
▲
by
Asm2D
7y ago
asmjit has a register allocator, for sure not the highest quality one, but it's there in the asmjit's Compiler infrastructure.
45.
▲
by
Asm2D
7y ago
Thanks! Let's continue on that issue page.
46.
▲
by
Asm2D
7y ago
Yeah there is definitely some redundancy here, but there were reasons I wrote these tools. There are multiple sources that provide instruction timings. Agner's instruction tables are amazing, but they cover only a single microarchitect
47.
▲
by
Asm2D
7y ago
Thanks for the article! I started AsmGrid project that provides X86 architecture overview and instruction timings for people like me that often work with assembly. It's basically a simple web application that displays data provided by
48.
▲
by
Asm2D
8y ago
I don't know what rasterizers you refer to with: "CPU rasterization is always O(N) complex (where N is number of pixels on screen)" But this is definitely not the Blend2D case. I think you will not find rasterizers in
49.
▲
by
Asm2D
8y ago
It's true that increasing the size of framebuffer demands more from CPU as well. According to my experience a single core on a modern machine has no problem to render real-time into a FullHD framebuffer at high frame rate (depending on
50.
▲
by
Asm2D
8y ago
Really appreciated, thank you!
51.
▲
by
Asm2D
8y ago
Blend2D uses dense cell-buffer similarly to font-rs, however, it works quite differently and this difference allows Blend2D to be efficient even when rendering large paths: - Dense cell buffer, 32-bit integer per one cell (FreeType/Qt