3 ms·
Outside of a benchmark- that's latency you can hide though, right? Without reference to the specifics of this situation, which I don't claim to know much about,
by barrettcolin 16y ago
Outside of a benchmark- that's latency you can hide though, right? Without reference to the specifics of this situation, which I don't claim to know much about, the co-processor can be sorting one dataset in it's local memory while the CPU is transferring data into another local buffer (or even, the co-processor can initiate the transfer itself using DMA).
- whakojacko 16y agoOnly to a limited extent. Assuming this is a regular-ish implementation of radix sort, you could start doing the first bucketing while data is being copied in. Likewise, you can copy out the sections which already have been fully sorted. But I would be surprised if you saw more than 10-15% percent speedup over naive copy execute copy.
- barrettcolin 16y agoYour observation is likely correct, but I was talking about sorting two or more independent sets, not partitioning the sort of one set; assuming that a "proper" (non benchmark) application will have several workloads waiting to be processed at a given time, so the GPU can be working on one while the data for the next is being transferred (so hiding the latency for the transfer for all but the first set). [edit] I don't know how practical that is on real world hardware; a billion keys is obviously a lot of data to transfer.