Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
celrod
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
27 ms
·
241.
▲
by
celrod
6y ago
A few of the issues: 1) Not much code is written to take advantage of them. This seems like inherently less of a problem with SVE, assuming it becomes widely adopted, as this should naturally support compiling source to take advantage of wh
242.
▲
by
celrod
6y ago
OpenBLAS supports dispatching based on architecture. On Skylake-X, it was using SKYLAKEX kernels, not NEHALEM.
243.
▲
by
celrod
6y ago
I think this was a great comment and discussion thread on OOO vs VLIW: https://news.ycombinator.com/item?id=24466593 The biggest takeaway for me was the importance of scalability. As OOO CPUs get better, they can become inc
244.
▲
by
celrod
6y ago
If I understood lend000 correctly: > The fact that Intel's only new releases on 10nm are <=6 core low power products suggests their yields are still poor. Their argument was that, had they followed AMD's approach, they would
245.
▲
by
celrod
6y ago
> Use subtyping in Julia and enums in Rust. I would use enums in Julia, making the enum you define a field of the PlayerClass: @enum PlayerStar Sol Pol Cent struct PlayerClass star::PlayerStar # other fields end Ju
246.
▲
by
celrod
6y ago
Hmm. FWIW, on my [Skylake/Cascadelake]-X Intel systems, Intel's compilers performed well, almost always outperforming GCC and Clang. But on Zen, their performance was terrible. So I was happy to see that MKL, unlike the compilers,
247.
▲
by
celrod
6y ago
No, I don't. The 32-core AWS systems must be Epyc, so I'll try benchmarking there. When OpenBLAS identifies the arch, it is competitive with MKL in single threaded performance, at least for matrices with a couple hundred rows and
248.
▲
by
celrod
6y ago
I think MKL actually fixed Zen performance. That is, the workaround no longer makes any difference because it is no longer needed. Small matrix multiply benchmarks on a Zen2 (Ryzen 7 4700U), featuring MKL 2020.1.216+0, OpenBLAS, and Eigen:
249.
▲
by
celrod
6y ago
1. Dependency graph for cells, letting it automatically rerun what's needed when you change one. This keeps everything up to date. 2. git-friendly.
250.
▲
by
celrod
6y ago
> the x86_64 patents should expire this year. Apple could have probably made their own x86 chip without paying royalties. The free/open architecture that we've always dreamed of, might turn out to be the one we're already
251.
▲
by
celrod
6y ago
Note that HEDT Skylake-X all let you choose the clock speed for each license in the bios. It is definitely not a "we won't let you".
252.
▲
by
celrod
6y ago
I should include more detailed instructions. Dependencies include: eigen, gcc, clang, and (harder) the Intel compilers, as well as a recent version of Julia. If you have those, you should be able to instantiate the Manifest: https:/&#
253.
▲
by
celrod
6y ago
For the limited scope of rectangular loop nests without loop carried dependencies, my Julia library LoopVectorization.jl is able to vectorize and achieve often significantly better performance than Clang, GCC, or the Intel compilers; see th
254.
▲
by
celrod
6y ago
Julia does at least inline fairly aggressively, and it respects the `@inline` macro so long as the call is type stable. In practice, it may often inline more aggressively as there aren't any boundaries that will prevent inlining. If yo
255.
▲
by
celrod
6y ago
Thanks for sharing the benchmarks. I made the arrays huge to get roughly the same ball park of times you reported. I also preallocated the memory in that benchmark to avoid unnecessary allocations (the observed allocations are because multi
256.
▲
by
celrod
6y ago
I tried element-wise sum, and got julia> using LoopVectorization, BenchmarkTools julia> A = rand(10_000, 10_000); B = rand(10_000, 10_000); C = similar(A); julia> @benchmark vmapntt!(+, $C, $A, $B) BenchmarkTools.Trial
257.
▲
by
celrod
6y ago
I don't really want to sacrifice runtime performance for compile time performance, so long as the compilation time is "reasonable", but so far we've mostly been having our cake and eating it too. The eliminating-invalida
258.
▲
by
celrod
6y ago
Yes, that is why it is "[s]uggesting getting a cat is a good idea for someone concerned about depression"
259.
▲
by
celrod
6y ago
Note that the article supports your point: the number of cats at home had a negative effect on depression (p = 0.021). Suggesting that getting a cat is a good idea for someone concerned about depression. Of course, it may be that depressed
260.
▲
by
celrod
6y ago
Ping pong is also interesting. Expensive paddles have sticky rubber that can help experienced players control the spin on a ball. Cheap paddles don't have much grip. Very cheap paddles, the hard ones with short pips and no sponge, have
261.
▲
by
celrod
6y ago
They got an even bigger wrecker (with a third axle, helping to spread the weight). The end of the article said it freed the home, Big Blue, one yellow wrecker, and was then working on the last.
262.
▲
by
celrod
6y ago
My dad passed away late last year. He was a member of the 200 mph club, and still holds a few records he set in the 70s/80s. Perhaps my favorite story was of the motorcycle he used to break 200mph on the saltflats. Some called it the h
263.
▲
by
celrod
6y ago
> If fatigue is a simple threshold (i.e. you "become fatigued" after a certain number of units of work done), then the obvious solution is to work the employees at 100% utilization for half as long (at which point they should b
264.
▲
by
celrod
6y ago
For gcc you need `-funroll-loops` to unroll and `-fvariable-expansion-in-unroller` to get multiple accumulation vectors. By default, it'll only use 2 accumulation vectors. You can set it to 4 (for example) with `--param max-variable-ex
265.
▲
by
celrod
6y ago
For BLAS in particular, this paper can give you an idea of some of MLIR's capabilities: https://arxiv.org/pdf/2003.00532.pdf (But maybe you already know them better than I do.) LoopVectorization can't do many
266.
▲
by
celrod
6y ago
That's what I did when calling OpenBLAS and MKL, but I confess I don't know the internal details of a non-inlined `matmul` call in gfortran when you don't use `-fexternal-blas`. Just writing three loops and letting the compil
267.
▲
by
celrod
6y ago
There is no reason it could not. Those optimizations just have to be implemented. Flang (the one merged into LLVM) is using MLIR, which has all the required code-gen abilities. That just leaves the cost modeling / deciding which optimi
268.
▲
by
celrod
6y ago
My color coding is bad (and there's an open issue about it), but I haven't come up with a great solution. One idea I should add was to make everything relative to LoopVectorization, but that wouldn't help for those interested
269.
▲
by
celrod
6y ago
I think Eigen's template strategy is probably hard to beat from the perspective of combining operations. How well Fortran does is probably mostly up to the compiler implementation. In some of my benchmarks, gfortran sometimes seems to
270.
▲
by
celrod
6y ago
https://travisdowns.github.io/blog/2020/01/17/avxfreq1.html AVX (256 bit) instructions also suffer a penalty. Both the 256 and 512 bit instructions resulted in a slowdown for 9 microseconds featuring a q
More ›