4 ms·
I wonder if the inclusion of NVLink in Power 8+ will cause Power to excel in ML applications. It could well be quite a bit faster than x86 just due to the memor
by virtuallynathan 11y ago
I wonder if the inclusion of NVLink in Power 8+ will cause Power to excel in ML applications. It could well be quite a bit faster than x86 just due to the memory/interconnect bandwidth.
- PeCaN 11y agoNVLink and CAPI[1] both have huge potential for machine learning. However, a lot of the benefits of NVLink for ML come from GPU-to-GPU NVLink, which doesn't require CPU support. 1. CAPI doesn't seem to get mentioned to much around here, but imagine an FPGA directly accessing some shared system memory. It's neat.
- DiabloD3 11y agoIt does seem to require support in the PCI-E host controller, which for both modern Intel and those POWER machines, is on die on the CPU. So, it "requires" CPU support, just not in the way usually meant.
- mikehollinger 11y agoCorrect; the CPU and the end-point accelerator both must cooperate to negotiate the CAPI link. Disclaimer: I work on this with some very smart people @ IBM. Opinions are my own. When a PCIe device is in CAPI mode, the PCIe protocol is used as a transport layer, but the CAPI protocol rides on top, and hardware in the CPU's PHB (the CAPP unit) and hardware in the accelerator (the PSL in this case) cooperate to present the common address space to the process and to the accelerator itself. [1] If a CAPI-capable card's plugged in to a non-CAPI-capable slot, it remains a PCI card. If a non-CAPI card's plugged in to a CAPI-capable system, it remains a PCI card. If both sides match on protocol versions and the kernel contains the cxl driver, the kernel will switch the slot into CAPI mode, and the CAPP unit and PSL effectively take over the PCI link on either side. [1] http://events.linuxfoundation.org/sites/events/files/slides/lcjp15_mackerras.pdf http://events.linuxfoundation.org/sites/events/files/slides/... - see page 13+ for some GPU / NVLink materials, and page 24+ for CAPI materials (oh - and I worked on the product who's data is quoted on page 29 [2]! [2] https://www.ibm.com/developerworks/community/blogs/fe313521-2e95-46f2-817d-44a4f27eba32/entry/power8_capi_flash_in_memory_expansion_to_speed_data_access?lang=en https://www.ibm.com/developerworks/community/blogs/fe313521-...
- bogomipz 11y agoOh is CAPI an onboard FPGA with a memory controller?
- ajdlinux 11y agoCAPI allows an FPGA connected via PCIe to be treated as a coherent peer to the CPU cores that is able to hold cache lines and also use address translation. Among other things, from the application programmer's perspective, the CAPI accelerator can basically be treated as if it were another thread, since it can use the application's virtual address space - the application can set up data structures in main memory and pass unmodified pointers to the CAPI card. http://www-304.ibm.com/webapp/set2/sas/f/capi/CAPI_POWER8.pdf http://www-304.ibm.com/webapp/set2/sas/f/capi/CAPI_POWER8.pd... is a good intro. [Disclosure: I work on CAPI at IBM]
- bogomipz 11y agoThanks for the link. The papers mentions key/value stores. Would a valid use case for CAPI be something similar to a "flash cache" where the FPGA is not as fast as DRAM but still faster than NAND flash?
- mikehollinger 11y agoYeah, it's neat. (I work on stuff that exploits this). We open-sourced the software side of our first flash IO accelerator last year. [1] You can do some pretty cool things from a HW designer's perspective inside the accelerator, and in the main application. Since the accelerator is cache-coherent, and able to map the same virtual addresses as a given process (and attach to multiple processes' address spaces) the device can do "simple" things like follow pointers, which used to require building a command / data packet, DMA'ing it to the device, and then waiting for a response packet. This, effectively, frees up the main CPU to do other things, rather than wrangle data. It also means that bottlenecks move. [1] https://github.com/open-power/capiflash https://github.com/open-power/capiflash
- bogomipz 11y agoSo the idea is to present nand a memory device rather than a block device?
- Symmetry 11y agoFor huge datasets the GPU-CPU links might become more important with Pascal now that the GPUs are allowed to trigger page faults.
- Symmetry 11y agoThinking about this some more no way would the latency of PCIe ever be as big as that of paging in some memory from disk. So page faults can't really make this more important in Pascal.
- bogomipz 11y agoCan you elaborate on CAPI/NVlink would be beneficial to an ML workload?
- PeCaN 11y agoNVLink is similar to a higher-bandwidth PCIe connection, except that multiple GPUs can be connected with it. It's primarily useful for very large convnets, which use a lot of memory and can be bandwidth-limited. It doesn't require any particular modifications to a model or framework to take advantage of it. CAPI is much more flexible and interesting. It allows a CAPI-capable connected device access to a process's virtual memory. Essentially, you can extend the CPU's capabilities with CAPI. Usually this would be an FPGA (and the utility of FPGAs for machine learning is very much a research topic), but I could easily see a DSP being useful for voice recognition. GPUs can take advantage of it too, but ML work is usually just offloaded entirely to the GPU. CAPI is very very cool and designed by some very smart people. I'm excited what people will do with CAPI and FPGAs.