3 ms·
Actually, the reason I didn't show distributed was in fact the naive original implementation of the TCP stack support. The platform minimizes data movement over
by jcbeard 10y ago
Actually, the reason I didn't show distributed was in fact the naive original implementation of the TCP stack support. The platform minimizes data movement over network pipes, however due to time constraints on me finishing a PhD I chose to get the more interesting mathematical modeling aspects worked out rather than focus on engineering something that is pretty well understood (especially using techniques directly developed for HPC with MPI). Will push better support soon.
On the multi-process..yes, that's easy. I've removed it for the current main-line branch given the lack of demand. IfDef'ing the code made it much easier to proceed with getting it ready for alpha. The FIFO mechanisms are well tested using SHM and the forking code will be added back in soon. Another thing I commented out to get it working on multiple platforms is the NUMA placement code, now that I think hwloc will work on all platforms I'll get it added back in. Helps out on cross-socket communication quite a bit, as well as placing buffers closest to PCIe root for data transfer to accelerators.
In reality the data movement is no worse than any OpenMP or other parallel program. In as many places as we can, the data is left in place vs. pushed. Between nodes, it gets more fun...however it's still a rather well understood problem. Thanks again for the interest! I'll see if I can do a ShowHN before CPPNow 2017 for the beta release.
- vmarsy 10y agoThanks for commenting! > Between nodes, it gets more fun...however it's still a rather well understood problem. Indeed it's understood, but the scalability problems switch somewhere else once your communication becomes the bottleneck. One advantage of dataflow-like paradigms like yours is that you can avoid major barriers like in Bulk-Syhnchronous-style programs. Those barriers start to become costly when you run with ten/hundreds of thousands processors. That'd be super interesting to see your library scale in those settings, similar to [1]. [1] https://www.researchgate.net/profile/Mani_Zandifar/publication/279286000_The_stapl_Skeleton_Framework/links/56327c9208ae242468d9f8df.pdf https://www.researchgate.net/profile/Mani_Zandifar/publicati...