Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
voidstarcpp
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
voidstarcpp
3y ago
>But if our performance knowledge is outdated, how do we know which cases to separate the code into? Case separation is less of a commitment than deciding on the specific optimization yourself. You know empirically which cases are used i
2.
▲
by
voidstarcpp
3y ago
The "tiny math function" is used for expository purposes of the generic application, to fit a trivial example on one page. Obviously it's not how you would literally write a function that transforms elements in a vector.
3.
▲
by
voidstarcpp
3y ago
>GCC is actually good at that, -O3 has no issue recognizing it. imo, if you have to go to O3 or enable a pragma to get an "obvious" optimization then this is undesirable and probably something the programmer still needs to be c
4.
▲
by
voidstarcpp
3y ago
Addressing the aliasing concern would be the easiest improvement. I observed in the assembly that the source pixel is being re-read all four times it is used, which could be fixed. Writing an optimal composite function is of course not real
5.
▲
by
voidstarcpp
3y ago
This is possible if the call site can see the implementation, but you can't count on it for separate translation units or larger functions. My goal was to not rely on site-specific optimization and instead have one separately compiled
6.
▲
by
voidstarcpp
3y ago
You need to use a compiler specific "always inline" directive if you want macro-like functionality of actually inlining code. On its own, the C++ "inline" keyword does not cause inlining to happen, although compilers may
7.
▲
Going faster by duplicating code
(voidstar.tech)
158 points
by
voidstarcpp
3y ago
|
68 comments
8.
▲
by
voidstarcpp
3y ago
TL;DR: If you copy paste the same implementation code in different branches, you give the compiler opportunities to generate faster code for each case it wouldn't have otherwise generated, without you having to do any manual optimizati
9.
▲
by
voidstarcpp
3y ago
In my limited experiments WAL wasn't faster for high-row-speed writes. When rows/s are maximized SQLite is CPU limited by internal bytecode operations rather than waiting for disk stuff.
10.
▲
by
voidstarcpp
3y ago
PostgreSQL supports heap tables, which should blow away SQLite's mandatory clustered index tables in an unindexed insert test.
11.
▲
by
voidstarcpp
3y ago
With SQLite by default nobody else can even read the database while it's being written, so your comment would be better directed to a conventional database server.
12.
▲
by
voidstarcpp
3y ago
>Parsing the very simple SQL doesn't seem like it would account for much time, so is the extra time spent redoing query planning or something else? If you're inserting one million rows, even 5 microseconds of parse and planning