Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
celrod
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
23 ms
·
211.
▲
by
celrod
6y ago
Ah. I have AMD GPUs on desktop, and integrated graphics on laptops.
212.
▲
by
celrod
6y ago
I've been using Linux almost exclusively for many years, and I had to google what screen tearing is after all the comments about it in this thread. I do get people's resistance to switch. I bought a M1 Mac Mini because I wanted to
213.
▲
by
celrod
6y ago
To tie Octavian.jl into this memory allocation discussion: Octavian uses stack-allocated temporaries when "packing" left matrix ("A" in "A*B"). These temporaries can have tens of thousands of elements, so that&
214.
▲
by
celrod
6y ago
> or maybe even custom compilers Like for autodiff or GPUs.
215.
▲
by
celrod
6y ago
Smallpox.
216.
▲
by
celrod
6y ago
I often end every line with a semicolon, so that it doesn't flood a REPL if I run it there. IIRC, groupby hasn't been optimized in DataFrames.jl yet.
217.
▲
by
celrod
6y ago
While moving the loop into Julia (as others suggested) is probably the better option, an alternative you could consider is DaemonMode: https://github.com/dmolina/DaemonMode.jl I.e., have a background Julia process so t
218.
▲
by
celrod
6y ago
4.6 GHz all core with AVX512 surprised me. At least AVX512 downclocking seems to be a Skylake problem. Ice Lake (which Rocket Lake is a backport of) similarly doesn't downclock much: https://travisdowns.github.io/blog&#
219.
▲
by
celrod
6y ago
Has ispc been deprecated? Last GitHub commit was three days ago.
220.
▲
by
celrod
6y ago
It is recreated on every call to `fib`, but the `_fib`s get to reuse it. Using a cache makes `_fib` O(N) to compute, so it gives a performance benefit as long as you aren't calling `fib(n)` with very small values of `n`.
221.
▲
by
celrod
6y ago
Cool, and fantastic summary! I enjoyed reading it.
222.
▲
by
celrod
6y ago
> (* Technically not all mutable objects live on the heap, because some never live at all, as they are optimized away so are never allocated in the first place.) The compiler will often stack allocate mutable objects in Julia. This is no
223.
▲
by
celrod
6y ago
Yeah, I'd like to know more about what it can do. Can it optimize sequences of shuffles? If so, it'd be worth a try, even if it takes hours (probably doesn't take that long?), I could let it run until it finds faster sequence
224.
▲
by
celrod
6y ago
I have on occasion, and when working on custom array types I've added support for being offset with optionally compile-time known offsets. As StrideArrays.jl matures and I write more libraries making use of it, I may use 0-based indice
225.
▲
by
celrod
6y ago
It's related to 0-based indexing in that if you you want to take/iterate over the first `N` elements, `0:N` works with 0-based indexing + close-open, but if you had 1-based and close-open, you'd need the awkward `1:N+1`. This
226.
▲
by
celrod
6y ago
I'm using gcc 10.2.0. I tried clang 11 and got more or less the same thing, so it doesn't seem to make much of a difference. Neither did messing with flags, like (I tried -fno-semantic-interposition -march=native and a few others)
227.
▲
by
celrod
6y ago
I just ran it again, and got more or less the same results: N = 1000000000, 953.7 MB starting experiments. two : 29.7 ns two+ : 36.5 ns three: 43.8 ns This surprises me. Normally, it does very well in most benchmarks I run.
228.
▲
by
celrod
6y ago
Yeah, that is strange. Why was it so slow? Trying on a desktop with a 7900X, I get N = 1000000000, 953.7 MB starting experiments. two : 17.7 ns two+ : 19.1 ns three: 26.4 ns bogus 1422321000 This is again close to 50% slo
229.
▲
by
celrod
6y ago
I have a Dell XPS 13 with a Tiger Lake CPU. Out of curiosity, running the script: > ./two_or_three N = 1000000000, 953.7 MB starting experiments. two : 30.6 ns two+ : 39.6 ns three: 45.1 ns bogus 1422321000 This
230.
▲
by
celrod
6y ago
They all died, male and female alike, but a higher percent of the males left behind fossils. Meaning getting yourself killed is only part of the equation. The other part is that the manner in which you get yourself killed has to increase th
231.
▲
by
celrod
6y ago
It's used a lot for things like analyzing clinical trials, e.g making futility or early stopping calls in interims, or for meta analysis. JAGS may still be the most popular, at least in some companies, but Stan is starting to catch on
232.
▲
by
celrod
6y ago
Source for Zen3 having 4x per cycle? That's impressive and I'd like to read more about it. Another comment mentioned 4, 2 of which are fma. Works this imply that a matmul kernel for zen3 should have a 2:1 ratio of fmas to mul+add?
233.
▲
by
celrod
6y ago
Running on Rosetta (and looking like an x86 cpu without fma) is likely a contributing factor, but the M1 was much slower for sums and sums of squares than a Cascadelake (Intel, avx512) cpu here: https://github.com/chriselrod
234.
▲
by
celrod
6y ago
Just to clarify, those precompiled caches are (currently) not of binaries. At the very least, LLVM still needs to generate code after loading, and often everything up to inference still needs to be run as well. Julia 1.6 involved a lot of w
235.
▲
by
celrod
6y ago
Also, oneAPI.jl: https://github.com/JuliaGPU/oneAPI.jl
236.
▲
by
celrod
6y ago
ISPC can be really good at SIMD-ing complicated control flow (ray tracers being the archetypal example). I'm interested in eventually working on something like that for Julia. In the mean time, it should be possible to deliberately wri
237.
▲
by
celrod
6y ago
Definitely. Once the AVX512-IFMA instruction set becomes more common (only available on Intel 10nm CPUs so far), I may try a PCG with 104 bits of state. Wanting SIMD rngs (separate generator per vector lane), PCG is hurt by 64-bit integer m
238.
▲
by
celrod
6y ago
Xoshiro over PCG arguments: http://pcg.di.unimi.it/pcg.php Response by the author: https://www.pcg-random.org/posts/on-vignas-pcg-critique.html Personally, I use a SIMD Xoshiro256+ when I want a solid
239.
▲
by
celrod
6y ago
I've ordered a tiger lake laptop which should arrive around the end of the month, so I'll be able to test then. So just speculation for now: I'd think AVX512 would still be advantageous for gemm kernels. If you use a 16 x 14
240.
▲
by
celrod
6y ago
Gaming doesn't really have SIMD or FP at all though, does it?
More ›