4 ms·
Yes, you can get good performance on GPUs due to the parallelism and memory bandwidth. On a new GPU like H100, I believe you can do ~ 50-100 GTEPS (billions of
by pgera 3y ago
Yes, you can get good performance on GPUs due to the parallelism and memory bandwidth. On a new GPU like H100, I believe you can do ~ 50-100 GTEPS (billions of traversed edges per sec) in a BFS. I'm not sure where the state of the art on CPUs is at, but you can certainly do efficient implementations on CPUs too. This paper is a few years old (https://people.csail.mit.edu/jshun/spaa2018.pdf https://people.csail.mit.edu/jshun/spaa2018.pdf), but has some numbers.
For CPU based index compression/decompression, the webgraph framework is quite mature and widely used and there's also Ligra+ that does it.
- WinLychee 3y agoVery cool, will digest all this! For my use-case CPU + lots of RAM has been fast enough, and there's a balance between per-thread latency and throughput (serving data concurrently). I'm definitely interested in compressing down the graph data further to see if I can drop latency further, will see if I can adapt some of this. Super neat to see this on GPUs, will check that out too. Also a fan of https://github.com/frankmcsherry/COST https://github.com/frankmcsherry/COST if you've seen it before as well!