3 ms·
Your observation is likely correct, but I was talking about sorting two or more independent sets, not partitioning the sort of one set; assuming that a "proper"
by barrettcolin 16y ago
Your observation is likely correct, but I was talking about sorting two or more independent sets, not partitioning the sort of one set; assuming that a "proper" (non benchmark) application will have several workloads waiting to be processed at a given time, so the GPU can be working on one while the data for the next is being transferred (so hiding the latency for the transfer for all but the first set).
[edit] I don't know how practical that is on real world hardware; a billion keys is obviously a lot of data to transfer.