Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
xoranth
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
61.
▲
by
xoranth
3y ago
(Sorry, I was wrong. See here for the author's reply: https://news.ycombinator.com/item?id=36627385 )
62.
▲
by
xoranth
3y ago
> I guess the compiler's unrolling heuristics generally aren't as good as that blocking "mod then div" alternative to Duff's device? Obviously taking `s` out of the loop condition is part of the magic. The magic
63.
▲
by
xoranth
3y ago
Interesting. I think you can vectorize the prologue using movemask + popcnt instead of keeping a counter in the ymm registers (warning: untested code, still need to benchmark it): const __m256i zero = _mm256_setzero_si256(); const
64.
▲
by
xoranth
3y ago
I rerun the benchmark vs loop-5 and loop-7 from the second post. Runtime is basically the same on my machine. I would have expected yours to be faster given that it needs to execute fewer instructions per loop iteration. Though maybe the CP
65.
▲
by
xoranth
3y ago
I think I managed to improve on both this post, and its sequel, at the cost of specializing the function for the case of a string made only of 's' and 'p'. The benchmark only tests strings made of 's' and '
66.
▲
by
xoranth
3y ago
Most blogs that have RSS also have a `<link rel="alternate" type="application/rss+xml">` tag that redirects you to the RSS feed. If you pass the link to the homepage to a feed reader[^0], it will follow the li
67.
▲
by
xoranth
3y ago
https://xoranth.net/ In-depth blogging about low level optimization. Two posts so far: - https://xoranth.net/memcmp-avx2/ A walkthrough of an highly optimized implementation of string comparisons. - h
68.
▲
by
xoranth
3y ago
Lower fees. Also, "exchange competition" is a thing in US stocks, but not, for example, in the futures' market. And there's HFTs in futures too, so having a single exchange wouldn't "remove" HFTs (not that
69.
▲
by
xoranth
3y ago
Your vanilla Debian box uses BMI2 whenever you do a string comparison, unless you are on a decade old CPU [^0]. The "strange" instructions are actually not that niche, it's just that usage tends to be "indirect" and
70.
▲
by
xoranth
3y ago
> Another simple trick is duplicating the buffer pointer field and buffer size field for consumer and producer Do you mean something like this? struct Inner { int\* data; size_t size; size_t cachedIdx; }
71.
▲
by
xoranth
3y ago
From the slides on `conflict`/`vpcompress`, it says. > In fact, we can re-use an existing structure with some minor changes What is the existing structure that they reused? The slides don't seem to specify.
72.
▲
Optimizing the `pext` perfect hash function
(xoranth.net)
3 points
by
xoranth
3y ago
|
0 comments
73.
▲
by
xoranth
3y ago
As a datapoint, numpy also chooses between radix and timsort based on data type. From the docs [^0]: > ‘stable’ automatically chooses the best stable sorting algorithm for the data type being sorted. It, along with ‘mergesort’ is current
74.
▲
by
xoranth
3y ago
They have sample searches for "best headphones", "steve jobs", and "python exceptions" on their main website [^0]. [0] https://kagi.com/
75.
▲
by
xoranth
3y ago
> For larger arrays, I believe the branchy search algorithms start to do better since half their branches get predicted correctly. (not sure if this is what you meant by "conflict misses") I believe the reason branchy binary se
76.
▲
A look inside memcmp on Intel AVX2 hardware
(xoranth.net)
3 points
by
xoranth
3y ago
|
0 comments
77.
▲
by
xoranth
3y ago
Thanks!
78.
▲
Ask HN: Resources on Memory Prefetching
2 points
by
xoranth
3y ago
|
2 comments
79.
▲
by
xoranth
3y ago
Thank you! I'll take a look.
80.
▲
by
xoranth
3y ago
> 2. glibc strings functions are ok, but not particularly good, ime Do you mean for RISC-V, or in general? (and in particular, what about x64?) What problems do they have? [Also what are in your experience better implementations?]
81.
▲
by
xoranth
3y ago
Mmm, it looks like LLVM even has a pass to actively convert generated cmovs into branches if one of the operands touches memory [^1], and that only supported way to force things to be brancheless is inline assembly [^2]. [^1] https:/&
82.
▲
by
xoranth
3y ago
> However, I’ve found that some compilers, e.g. GCC on x86-64 will refuse to make this variant branchless. I hate how fickle compilers can be sometimes, and I wish compilers exposed not just the likely/unlikely attributes, but also
83.
▲
by
xoranth
8y ago
I cannot speak on Carreyrou, but parts of _When Genius Failed_ make me think that Lowenstein did not have a good grasp of the subject he was writing about. For example, when discussing LTCM volatility strategies: >The stock market, for i
84.
▲
by
xoranth
10y ago
Could you share more details about the robust regression you are using? All resources I could find online on robust regression would either point to Laplace distributed residuals, or some capped loss function.
85.
▲
by
xoranth
10y ago
The difference in the drop between UKX and IAS is due to fx. Namely, if you were a British investor that bought an european index ETF before the move, you would have lost on the drop, but would have made some of the money back on the fact t