4 ms·
Been working on the API improvements for this for quite a while now. If there are any questions I can answer, I'll tell you whatever I can. (edit: in case it's
by tmurray 16y ago
Been working on the API improvements for this for quite a while now. If there are any questions I can answer, I'll tell you whatever I can.
(edit: in case it's not obvious, I work for NVIDIA on CUDA)
- deleted 16y ago[deleted]
- lrm242 16y agoI haven't done any CUDA programming, however I have wanted to dive in. I've been to some conferences where the breakpoint between whether CUDA gives enough benefit to the computation is largely driven by time to load a dataset into memory on the GPU. From what I've been told this has limited the use of CUDA to batch processing on large scale datasets. Is this still a limitation?
- tmurray 16y agoPCIe can be a limitation, but there are a lot of ways to amortize the latency of copying data to/from the GPU. Overlapping transfers with kernels, direct load/store from system memory to a kernel, multiple kernels running on the chip at the same time--there are a lot of things you can do. But in general, you're looking at a runtime of at least a few hundred microseconds before you're going to be able to get a benefit from the GPU.
- marshray 16y agoCongrats on your upcoming major release! I'm sure it was a lot of work. I'm a detail freak so I worry about this kind of thing more than most, but wouldn't a unified address space tend to hide large bus transfers from the programmer? Will the APIs still let you anticipate where the costs are, or are developers dependent on Windows-only profiling tools?
- tmurray 16y agoIt's not a true unified address space in the way I think you're getting at. We carve out the address space, but you can't just directly dereference a GPU pointer from CPU code or vice-versa. What we guarantee is that you can determine the location of a pointer based only on the value of the pointer so you can make the appropriate API calls to do a copy if necessary.
- iskander 16y agoIf the programmer still has to perform manual copies, what's the advantage?
- tmurray 16y agoYou can take an arbitrary pointer value in a function, look up its location, and do the copy if necessary, instead of requiring multiple entry points for CPU and GPU pointers. This also simplifies multi-GPU programming a lot--copy from one arbitrary pointer to another and it just works, instead of having to specify which device the copies are going to or anything like that.
- iskander 16y agoOne thing that completely puzzles me: How does dynamic memory allocation/deallocation in GPU kernels work? Also, will the programmer be able to use virtual functions in CUDA kernels? If so, does the usual thinking about warp divergence apply? That is--- should all threads in a warp ideally be calling the same function address?
- tmurray 16y agoI think there's some preallocated memory for dynamic allocations from the device. As far as I know virtual functions work as well, and the usual thinking about warp divergence does apply.
- maximilian 16y agoHave you done any benchmarks regarding inter-GPU memory bandwidth (i.e. GPU1->GPU2) vs. GPU1->host->GPU2? I am working on multi-gpu problems, which are often throttled by bandwidth limitations. Does the new SDK have examples of this and the new unified memory architecture (which looks handy)?