Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
poulson
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
poulson
13y ago
Interesting. That isn't excessively large for distributed-memory Krylov-subspace eigensolvers (e.g., SLEPc). Do you have a good reference for the current algorithms used? And how large is the nullspace typically? EDIT: I assume this [1
32.
▲
by
poulson
13y ago
Would you be willing to answer what sort of linear algebra is required in the function field sieve that the talk claims is difficult to parallelize? I write open source parallel linear algebra software [1] and would be happy to contribute.
33.
▲
by
poulson
13y ago
Yes, I should have mentioned this. SLEPc has the only distributed implementation of partial reorthogonalization (the key component of high-performance Krylov SVD and Hermitian eigensolvers) that I'm aware of.
34.
▲
by
poulson
13y ago
Sparse-direct solvers use graph theory to exploit nonzeros. In particular, Clique uses recursive nodal bisection ("nested dissection") for this purpose. The separators from this process end up (more-or-less) becoming cliques in th
35.
▲
by
poulson
13y ago
Thanks! I actually work on a lot of fast/sparse linear algebra. For example, see Clique: http://github.com/poulson/Clique
36.
▲
by
poulson
13y ago
I appreciate the complements, but I disagree with a few of your points: 1. What granularity to distribute the entries in the matrix is a long and subtle argument which doesn't provide a clear winner for every operation (the current con
37.
▲
by
poulson
13y ago
The research behind ScaLAPACK was very worthwhile and led to a huge number of algorithms and insights, and it took me several years of earnest weekend/late-night work to get Elemental to its current state (often drawing from the previo
38.
▲
by
poulson
13y ago
There's no need to be a Negative Nancy, especially when you're affiliated with the project you're promoting (without disclosing it).
39.
▲
by
poulson
13y ago
Elemental implements , as opposed to wraps, the distributed-memory algorithms. In other words, it does not call any library like ScaLAPACK; it is an alternative approach. Elemental builds on top of BLAS/LAPACK/MPI in order to pro
40.
▲
by
poulson
13y ago
Essentially all distributed dense linear algebra libraries are built on top of sequential dense linear algebra libraries (mine as well, but on the interface rather than the implementation). Distributed libraries are at least an order of mag
41.
▲
by
poulson
13y ago
With all due respect, did you even read the title of the post? This is a distributed-memory library, unlike all of the ones you just mentioned. This is a fundamental difference in design and capability. The only related libraries are ScaLAP
42.
▲
by
poulson
13y ago
Agreed. Accelerators can certainly lead to speedups, but it's hard to overemphasize how important it is to also have an optimized vanilla pure-MPI implementation (if nothing else, to compare the 'accelerated' version to).
43.
▲
by
poulson
13y ago
Accelerator support is something that happens within the node and is in some sense orthogonal to the high-level design. I recently received funding to add such support to the library (and I hope to add it within the next year). Also, not al
44.
▲
by
poulson
13y ago
Honestly, I doubt it. Getting added to such a list has both political and merit-based components.
45.
▲
by
poulson
13y ago
There is unfortunately no good answer to this question, as it would completely depend upon the algorithm, MPI implementation, communication network, vendor-tuned BLAS library, etc., etc. However, I can safely say that Elemental is always co
46.
▲
by
poulson
13y ago
I'm happy to answer any questions about the pros and cons of the library (and to remain quite objective). The main feature that the library currently lacks that is available in ScaLAPACK is a parallel Schur decomposition, but one is in
47.
▲
C++ Library for Linear Algebra on Supercomputers
(libelemental.org)
39 points
by
poulson
13y ago
|
38 comments