Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rohany
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
rohany
1y ago
> I always assumed that when one warp waits for results from a long latency instruction, another warp, potentially from another block can be scheduled in. Yes, that is correct. However, most MMA-style kernels that utilize the Tensor Core
2.
▲
by
rohany
1y ago
Author here! I think that warp specialization is inherently related to multi-stage pipelining, they aren't really alternatives of each other. Warp specialization is a way to realize a multi-stage pipeline in the face of hazards that ma
3.
▲
Unweaving warp specialization on modern tensor core GPUs
(rohany.github.io)
34 points
by
rohany
1y ago
|
4 comments
4.
▲
by
rohany
3y ago
Legion has been used to develop distributed, drop-in replacements for libraries like NumPy and SciPy Sparse -- https://github.com/nv-legate/cunumeric/ , https://github.com/nv-legate/legate.spar
5.
▲
by
rohany
3y ago
I agree! Much of this work was done as part of the overarching TACO project ( https://github.com/tensor-compiler/taco ), in an attempt to distribute sparse tensor computations ( https://rohany.github.io/pu
6.
▲
by
rohany
3y ago
I am the author of this paper -- kind of crazy to see it posted here on its own! I'll hang around to try and answer questions
7.
▲
by
rohany
6y ago
Enums should be coming in 20.2 as well :)
8.
▲
How Online Primary Key Changes Work in CockroachDB
(cockroachlabs.com)
22 points
by
rohany
6y ago
|
0 comments