11 ms·
Hi! I'm the primary author and maintainer of the library. Thanks for the interjection. The intended application for RaftLib is to make something that will scale
by jcbeard 10y ago
Hi! I'm the primary author and maintainer of the library. Thanks for the interjection. The intended application for RaftLib is to make something that will scale, just as you mention. I wrote this post a long time ago when I was just trying to get people interested in using it. It's a simple example that shows you can take many lines of standard parallel code, and write a much easier to read (smaller) version in a very short time that performs just as well or better than the manually managed parallel code.
A long time ago I was a biologist, then bioinformaticist. I wrote some code that would scale to a single node, and to a few dozen cores. In doing so, I realized how much I hated writing the same boilerplate code over and over again. TBB, c++11 threads, OpenMP, and MPI all basically have the same level of boilerplate and gotchas. I wanted to make something that was relatively easy to use and easily integrable with C/C++ code. Go was the only thing that came close, but it was brand new at the time.
It occurred to me while working on the AutoPipe system as a grad student that I could do something even better than a simple coordination language and at the same time subsume the functionality of a lot of parallel libraries. With stream/data-flow processing, I can do the exact same things I can do with OpenMP and MPI, but I can do more. The state encapsulation allows a whole host of cool optimizations, like identifying bottlenecks and duplicating actors dynamically (there's a whole host of reasons we'd be limited in OpenMP, c++11 threads). You can also compile an encapsulated function to another hardware platform entirely, or use high level synthesis tools to go to an FPGA (I'll be going there again soon too with RaftLib). The only thing that has to be constant across optimizations, is the connectivity of the DAG. By maintaining a port interface, just like you would hardware components (see Arvind's work from MIT...he's famous enough I just have to say Arvind :), we can compose really complicated applications. The port interface, it turns out, is also perfect for distributed compute.
Awhile back, I also had the realization that iostreams were perfect for this paradigm. Once you get your head around the concept, it seems quite natural. If it doesn't take off as a library, oh well. I enjoy working on it, and using it so I'll likely keep developing it in my spare time.
In the interim, I'll get back to exascale hardware stuffs :).
- Drdrdrq 10y agoIf I read your charts right, your app's single core performance was much better than pbzip2's, which is quite surprising. I thought these apps were severely optimized... Any comment?
- jcbeard 10y agoYup, it was quite a bit better on the upper end especially. Looking at the snoops on the bus using PAPI RaftLib does a better job at keeping the cache lines from bouncing. The benchmarked version also has a dynamically resizing FIFO which uses utilization of the queue itself to guide the sizing. This means that the FIFO can better adapt to dynamic behavior found in most applications run on top of an operating system (most all these days outside of HPC). Looking at load stalls, the RaftLib version has fewer, but not quite enough to account for the results. If you look at the single worker thread case, then jump to two threads..you can see a fairly big jump. RaftLib by definition is a pipelined programming system. The read file and compress are done perfectly in parallel. The bzip2 code doesn't quite pull it off in a perfectly pipelined fashion. It's close, but not quite. This results in less overlap of execution and communication. If I'd run on Linux (thread affinity on OS X is well, fun last time I checked..if not impossible to do manually), I'd also add thread affinity to the list which most people don't bother to optimize. Hot caches and synergistic cache accesses are quite beneficial.
- scott_s 10y agoI work on a streaming system (IBM Streams). I was reading through one of your blog posts (http://www.jonathanbeard.io/blog/2015/09/19/streaming-and-dataflow.html http://www.jonathanbeard.io/blog/2015/09/19/streaming-and-da...), and I saw this comment: 'Storm, Samza, and Spark are open-source streaming platforms that are focused on message-based data processing. These systems differ from other classic "streaming" modalities in that they eschew point to point communication for centralized brokers to distribute data.' That's maybe true for Spark, but it's not true for Storm and Samza (or of Streams). Yes, there is some centralization for application control, but there is no centralized data broker. That is, all messages are not routed through a centralized broker. I believe Storm does route all messages that come through the same JVM through the same logic and connection, but that's quite different from having a completely centralized broker. I can't speak for certain Samza, but I would be very surprised if it used a centralized broker. In Streams, I can say for certain that there is no centralized broker; messages are definitely sent "point to point."