Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
obl
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
17 ms
·
121.
▲
by
obl
8y ago
OpenGL is just an API, the underlying features are provided by the graphic driver. You can't ship with your own OpenGL, that would mean shipping with your own amd/nvidia/intel driver (and the associated kernel module, etc). A
122.
▲
by
obl
8y ago
you have to get the PR in the oven while it's still hot
123.
▲
by
obl
8y ago
Tangentially related, but maybe microsoft could actually do that : why is github's search so terribly bad ? If they had developed a good powerful code search (custom semantic engine for most used languages, complex queries, exact/
124.
▲
by
obl
8y ago
yeah. even LLVM is very slow for a JIT. Ask any project using LLVM as a backend for a JIT and they'll tell you it's an recurring issue. See for example https://webkit.org/blog/5852/introducing-the-b3-jit-
125.
▲
Specifying Memory Models Using Instantaneous Instruction Execution
(arxiv.org)
2 points
by
obl
8y ago
|
0 comments
126.
▲
by
obl
8y ago
very good point on the "addresses compare == iff same object" rule. In that case though, I think clang is right to optimize the callee (but it does introduce a problem in the caller) : the only place you could do the equality chec
127.
▲
by
obl
8y ago
Thanks. Unless something is escaping me, that's an optimizer bug. I'm pretty sure the ABI allows you to do whatever you want with the sret pointer, including passing it to another function to chain return for free. I guess you cou
128.
▲
by
obl
8y ago
> Note also that C doesn't have return-value-optimization, hence all your struct-returning functions will cause a call to memcpy (won't happen when compiled in C++ mode of course). What ? RVO is precisely needed because a copy
129.
▲
by
obl
8y ago
when the actual problem you are solving is uninteresting, you can at least come up with complex configuration puzzles and bugs to entertain yourself
130.
▲
by
obl
8y ago
HDLs are fine. I've found the tooling around them to be quite atrocious (slow, buggy, opaque, gui-based) though. Be it for synthesis, simulation, or even just compiler errors it's pretty bad compared to software.
131.
▲
ThinkPad A485
(www3.lenovo.com)
6 points
by
obl
8y ago
|
0 comments
132.
▲
by
obl
8y ago
in itself pointer chasing does not mean misusing CPU caches (your dataset could fit in cache and you could be pointer chasing inside it). the main issue that is strictly due to pointer chasing is that you risk having long dependency chains
133.
▲
by
obl
8y ago
right. his (1995) phd thesis is a classic for anyone interested in optimizing compilers http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.17....
134.
▲
by
obl
8y ago
The point of dynamic optimizations (such as ooo) is not only to hide implementation details (such as register file size) but very much to take advantage of dynamic opportunities that simply cannot be known statically. The optimal schedule c
135.
▲
Destination-Driven Code Generation [pdf]
(pdfs.semanticscholar.org)
2 points
by
obl
8y ago
|
0 comments
136.
▲
by
obl
8y ago
inverse bilinear map is a good example where, even if there is an analytical solution, it's often more robust to just do a couple newton iterations until you get down to machine precision. added benefit is that it carries on in 3D (alt
137.
▲
by
obl
8y ago
very curious if anyone has info on what he will be working on if it's any specific chip family. mainline x86 ? GEN gpus ? phis ? something else ?
138.
▲
by
obl
8y ago
That's fine if you're writing assembly. Compilers are free to use UB from the standard to optimize and absolutely do not guarantee you to emit machine loads and store naively as specified in the C source. At least for clang I'
139.
▲
by
obl
8y ago
running a smaller, continuous, current through them. PWM is done because it's cheap to implement given a fixed power rail and a digital control : need a single/couple of power switches that can even be integrated on die for low-is
140.
▲
by
obl
8y ago
yes, this is a troubling statement. I'd imagine there are some computerized systems between the manual counting and the actual result, but still.
141.
▲
by
obl
8y ago
you can saturate memory bandwidth without SIMD, since you can issue at least 2 8-byte scalar loads per cycle. it does not leave much room for actual processing though
142.
▲
by
obl
9y ago
I'm assuming it's because they wanted existing js/asm.js JITs to easily accept wasm, and those only see structured control flow.
143.
▲
by
obl
9y ago
any plan to move beyond "allocate linearly and never free anything" memory management ? semantics look nice and coherent but unfortunately without some kind (automatic or manual) of memory management it's not usable as a gene
144.
▲
Closing Greenlight Today, Steam Direct Launches June 13
(steamcommunity.com)
1 points
by
obl
9y ago
|
0 comments
145.
▲
by
obl
10y ago
I don't think the article states that the improvement in performance isn't the same for intel CPUs. maybe they just optimized the game.
146.
▲
by
obl
10y ago
xbox
147.
▲
by
obl
10y ago
is it ? those things are trivial enough to be entirely bandwidth limited. total_amount is 4 byte, passenger_count is 1 and those are tightly packed in a column layout. streaming through that in 150ms is almost within the reach of a single n
148.
▲
by
obl
10y ago
Would appreciate a reference on that if you have it handy. Is this more of a whole program measured X% regression on some perf benchs, or specific target cases where the codegen is notably worse ?
149.
▲
by
obl
10y ago
Unfortunately that's only superficially true. You have to care about GC interop, unwind & debug dwarf data, implementing part of the ABI in the frontend, and of course the occasional bugs when you exercise code path that clang does
150.
▲
by
obl
10y ago
Text templating can't possibly be better than (quasi)quoting. > you're essentially just writing crystal code even more so in a quasiquote, since you don't have to worry about adding extraneous parenthesis "just in cas
More ›