Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
xoranth
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
xoranth
3y ago
The latest version can [^1], though anecdotally I've seen clang/LLVM being smarter about it. [^1]: https://c.godbolt.org/z/vP8edfen7
32.
▲
by
xoranth
3y ago
> So you have to follow a comparison operation with either a conditional branch or a conditional move (and earlier processors in the x86 family didn't have conditional moves). The x86 family has the `setCC` instructions [^1] that mo
33.
▲
by
xoranth
3y ago
Yeah, makes sense. That said, (G)GP asked what could be done in raw assembly but not with C + intrinsics. My point is that conditional moves are one of usecases badly supported by compilers, and that (may) require dropping to assembly.
34.
▲
by
xoranth
3y ago
More prosaically, getting compilers to generate branchless code reliably is difficult, there's no intrinsic for CMOV* and similars, and the builtins that should act as an hint don't work[1]. [1]: https://c.godbolt.org&
35.
▲
by
xoranth
3y ago
Hindenburg doesn't short in order to drive down the price. Hindenburg compiles and releases information they believe to be truthful about a company, which they shorted beforehand in order to profit from the information. What I am talki
36.
▲
by
xoranth
3y ago
If you mean shorting Masimo purely to drive down the price, that's market manipulation and is definitely illegal. E.g. https://www.investor.gov/introduction-investing/investing-ba...
37.
▲
by
xoranth
3y ago
Thank you!
38.
▲
by
xoranth
3y ago
Amazing, thank you! PS: where/what can I follow to know when it is published? Could you share your colleague github handle (I guess he'll have one :) )?
39.
▲
by
xoranth
3y ago
Do you have any paper/talk that gives more details about the "geometric XOR filter"? If not, is there any plan to publish something?
40.
▲
by
xoranth
3y ago
It's supposed to be a single expression. You can write multiple lines by encasing the expression with parentheses.
41.
▲
by
xoranth
3y ago
Hint: it says "committed my password", as in a "git commit" :)
42.
▲
by
xoranth
3y ago
It's not (just) the variety of topics, it is that HN has way more users (and therefore way more submissions). A submission on lobste.rs will stay on top of recent/newest for days, on HN it might be minutes.
43.
▲
by
xoranth
3y ago
It is slower since it actually calls `memcpy`, instead of doing a single load. E.g. https://godbolt.org/z/5xa33qbar
44.
▲
by
xoranth
3y ago
> I fail to see how that is better. Null-terminated strings are a terrible idea in the first place. TLDR: you get extra bytes in your small strings Long version: Copying from FBString code: struct MediumLarge { Char* data_;
45.
▲
by
xoranth
3y ago
An even better representation is the one from `folly::FBString`. There, the tag that indicates small string mode is not a particular bound on capacity, but the last byte of the struct that is set to zero. That way, the tag also acts as a nu
46.
▲
by
xoranth
3y ago
I think he meant this: https://github.com/hanatos/vkdt
47.
▲
by
xoranth
3y ago
A ring buffer of pointers to structs is friendly to gather instructions. That said, the documentation shows a graph of operations applied to each packet. I'd expect that to lead to a lot of "divergence", and therefore being n
48.
▲
by
xoranth
3y ago
> For example, in a typical event driven architecture where you have worker threads running each on separate core, there would be something to decide where the task is queued and usually the logic will take into account how busy particul
49.
▲
by
xoranth
3y ago
> You are essentially asking the operating system to do the scheduling for you. But the OS will never be able to do it perfectly as it has no knowledge of what your application is doing. From LMAX presentations, it looks like they want y
50.
▲
by
xoranth
3y ago
In general, what are the advantage of a thread-per-request model? Better load balancing between cores?
51.
▲
by
xoranth
3y ago
> Hash table lookup is absolutely bounded by a constant, but going from L1 to L2 to L3 to RAM with anywhere between 0 and 5 TLB misses is far from constant. Wait, don't the inner nodes of the page table contain the _physical_ addres
52.
▲
by
xoranth
3y ago
With the sorting solution, you can stream the result directly to a file _without_ holding the entire file in memory. So it can be better if you are memory constrained (e.g. if the files are bigger than what you can hold in memory), or if yo
53.
▲
by
xoranth
3y ago
Given that the depth of the page table tree is fixed at 3/4/5 depending on the architecture and kernel, I think most people would be confused if asked whether the average complexity of a hash table lookup is always really O(1). Ev
54.
▲
by
xoranth
3y ago
You can generalize to > Now, given two N log files we want to generate a list of ‘loyal customers’ that meet the criteria of: (a) they came on ALL days, and (b) they visited at least L unique pages. while keeping linear time complexity a
55.
▲
by
xoranth
3y ago
Thank you! > ...until they are slower, because GPU's latency hiding mechanism (with occupancy) hides load latencies very well, while CPU just stalls the pipeline on every cache miss for ungodly amounts of time... Is the GPU latency
56.
▲
by
xoranth
3y ago
So, in NVIDIA parlance, my Skylake laptop would have 128 "cuda cores"? 128 = 4 (physical cores) * 2 (hyperthreading) * 8 (AVX2 f32 lanes) * 2 (floating point ports per core)
57.
▲
Ask HN: Was any Starfighter postmortem ever published?
186 points
by
xoranth
3y ago
|
169 comments
58.
▲
by
xoranth
3y ago
US exchanges tend to either bust or price-adjust clearly erroneous trades retroactively. That is good for unsophisticated customers, but it creates a disincentive for market makers to provide liquidity when there is a fat finger event. That
59.
▲
by
xoranth
3y ago
Do you have any example of particle systems without fluid sim? (e.g. videos on youtube from old games, or names of old games that used them?)
60.
▲
by
xoranth
3y ago
That's likely the fastest way to do that without vectorization. But you'd need to upcast 's' to an uint64 (or at least an uint32). That means that vectorization would operate on 32/64 bit lanes. With vectorization,
More ›