Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
celrod
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
32 ms
·
271.
▲
by
celrod
6y ago
I load a few packages in my startup. `--startup=no` makes the difference between about 0.12s and 0.5s on the 1+1 example for me. This doesn't really bother me, but maybe I should try PackageCompiler or adding `using OhMyREPL, Revise, C
272.
▲
by
celrod
6y ago
CPUs can execute multiple instructions per clock cycle (recent x86_64 can do 4-5). However, instructions can take 1 or (commonly) several clock cycles to complete, before their results are available for any instructions depending on said re
273.
▲
by
celrod
7y ago
I've a lot more experience with Julia than any other language (and am a huge fan/am heavily invested). My #2 is R, which has a much more basic type system than Julia. So -- as I don't have experience with languages with much
274.
▲
by
celrod
7y ago
I agree. Criticizing someone for being biased is an ad hominem; dismissing their arguments because of it is committing the genetic fallacy. In the extreme, it reminds me of this blog post, about how knowing about biases can let some folks u
275.
▲
by
celrod
7y ago
For small, statically sized arrays, LoopVectorization + tripple loops is also much faster than MArrays. LoopVectorization doesn't support SArrays yet, because you can't get pointers to them. MArrays will be stack allocated if they
276.
▲
by
celrod
7y ago
Interesting. Would this be safer in a language like Fortran, where (without aliasing between separate arrays), loop dependencies should be more obvious? I think it'd be nice to be able to activate this mode through pragmas. Does "
277.
▲
by
celrod
7y ago
I think that is possible. I've been thinking a lot about vectorizing loops recently, especially on AVX512 systems. I've mostly been doing microbenchmarks, and I realize that microbenchmarks might not give a realistic full-program
278.
▲
by
celrod
7y ago
I've been working on LoopVectorization in Julia, and benchmarking against a few compilers. Intel's compilers are far ahead of GCC and LLVM in vectorizing loops. LLVM (even with Polly) fairs worst in my benchmarks, so it does not l
279.
▲
by
celrod
7y ago
Correction: gcc 6 and 7 create them too. They just have 7 unrolled calls to log finite before (and then again after) a loop surrounding the vectorized call.
280.
▲
by
celrod
7y ago
Yes, here is the source for 8 double precision logs (with avx512) in glibc, for example: https://github.molgen.mpg.de/git-mirror/glibc/blob/20003c498... The "sysdeps/x86_64/fpu/multiarch&q
281.
▲
by
celrod
7y ago
Recent versions of GCC as well as the Intel compilers can autovectorize with the right flags (-ffast-math with GCC, -fast with Intel), and yield better performance with the <math.h> functions than you'll get from SLEEF on x86_64.
282.
▲
by
celrod
7y ago
Compare my username with that of the author of that library ;). For special functions, LoopVectorization relies on SLEEFPirates.jl, which is a fork of SLEEF.jl, a Julia port of version 2 of SLEEF. Most of the changes in SLEEFPirates are so
283.
▲
by
celrod
7y ago
I've been relying on a Julia port of version 2 (version 3 is out now). Version 3 added a lot of new functions, and I believe it improved performance on many of the old ones. It is much faster (when vectorized) than what you get in base
284.
▲
by
celrod
7y ago
On the desktop chips (x299) it's easy to adjust all the clock speeds in the bios. If the workloads I'm most interested in are all avx512-heavy (why I bought x299 instead of threadripper), do you think there'd be a reason to s
285.
▲
by
celrod
7y ago
OpenBLAS still didn't have adequate avx512 performance the last time I tried it (a few months ago). I think BLIS is the best library for performance portability. MKL offers a lot more than just BLAS and LAPACK. Would be cool for a proj
286.
▲
by
celrod
7y ago
Chrome has an app mode. From memory, the steps are roughly 1. Press the three dots to the right of the url. 2. Click more tools on the drop down menu. 3. Finally "create shortcut". This should let you open that website in a mode w
287.
▲
by
celrod
7y ago
f90 introduced recursion and allocatable (dynamically sized) arrays. So (in reference to tombert's comment), f90 and beyond doesn't have those limitations. It also introduced conveniences like operator overloading and element-wise
288.
▲
by
celrod
7y ago
Celsius S36. My GPU is on a separate loop and was idle at the time. I was running benchmarks of Intel MKL's zgemm vs zgemm3m because of a Julia PR that recommended replacing the former with the latter. I don't think anything hits
289.
▲
by
celrod
7y ago
Benchmarking a full sweep with 0 objects to free in Julia: julia> @benchmark GC.gc() BenchmarkTools.Trial: memory estimate: 0 bytes allocs estimate: 0 -------------- minimum time: 64.959 ms (100.00% GC) m
290.
▲
by
celrod
7y ago
The 7nm Ryzen parts should have double the avx2 as the older parts. Zen1 has half the throughput per cycle (or twice the reciprocal throughput) when using ymm (256 bit) registers vs xmm (128 bit) in general https://www.agner.org&
291.
▲
by
celrod
7y ago
You can look at binning statistics for non-avx/avx2/avx512 clock speeds: https://siliconlottery.com/pages/statistics For example, the worst 7980XEs do 4.1/3.8/3.6 GHz for each of these respectively.
292.
▲
by
celrod
7y ago
What do you think of Julia's macro-based approach? That is, there are `@inbounds` and `@fastmath` macros that turn off bounds checking/enable fast-math flags in the following expression. `@fastmath` works simply by swapping functi
293.
▲
by
celrod
7y ago
I am using Julia at Eli Lilly for fitting Bayesian models with MCMC.
294.
▲
by
celrod
7y ago
> I would still keep in mind that the metric worth using for benchmarking is the number of effective samples per second, and this also depends on the HMC variant you use. I was getting similar effective sample sizes/sample size in b
295.
▲
by
celrod
7y ago
I've been using Julia. I've been working on a front end meant to help specify vectorized models and their gradients. It is alpha-quality software (far from production ready), but here is the github: https://github.com&#
296.
▲
by
celrod
7y ago
Taking my 7980xe as an example: When it runs non-avx512 loads, I currently have it set to run at 4.1 GHz (all-core). When running avx-512 heavy loads, it instead runs at 3.6 GHz -- and tends to get much hotter (70-80C instead of 50-60C). 3.
297.
▲
by
celrod
7y ago
The wikipedia link shared above by aunty_helen ( https://en.wikipedia.org/wiki/Pavagada_Solar_Park ), which matches the name in the google maps link, says 53 km^2. It also says that it could produce 600MW by the end of
298.
▲
by
celrod
7y ago
This isn't about ego depletion, but about his mistakes on priming (which also failed to replicate), Kahneman writes: """ My position when I wrote “Thinking, Fast and Slow” was that if a large body of evidence published i
299.
▲
by
celrod
7y ago
On that note, it'd also be unwise to bet on the "ego depletion happens only to those who believe in it" finding from the paper I cited being replicate-able either -- it may be that the more broadly is no simple ego depletion
300.
▲
by
celrod
7y ago
People's beliefs about their responsibility and efficacy do have real impacts on their performance. Research has shown that whether or not someone's will power is limited is largely determined by whether or not they think it is. B
More ›