Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
pgera
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
pgera
3y ago
>You can also lay the graph array order out to minimize cache misses. This is something I've been looking into, but haven't implemented personally. The issue with RCM is that it only works for undirected graphs, from what I rem
2.
▲
by
pgera
3y ago
Yes, you can get good performance on GPUs due to the parallelism and memory bandwidth. On a new GPU like H100, I believe you can do ~ 50-100 GTEPS (billions of traversed edges per sec) in a BFS. I'm not sure where the state of the art
3.
▲
by
pgera
3y ago
> I believe you are not guaranteed for the edge data of adjacent nodes to be adjacent in memory The edge data of a particular node is contiguous, but yes, the edge data of a collection of nodes is not contiguous. You can reorder (permute
4.
▲
by
pgera
3y ago
Btw, core EF is quite efficient (perf wise) on the decoding side even on GPUs. I wanted to do PEF, but that seemed a bit more involved and I didn't have the time to do it. Here's a GPU implementation for graph problems if anyone i