3 ms·A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation20 points by matt_d 1mo ago