3 ms·
I found it curious the author mentioned using the Multiprocessing with pickle() but not Pipe(). Pickle() streams entire objects, while Pipe() can be used to se
by while_true_ 4y ago
I found it curious the author mentioned using the Multiprocessing with pickle() but not Pipe(). Pickle() streams entire objects, while Pipe() can be used to send data between processes. Maybe the latter is faster, especially if the data are short strings and the like?
- nijave 4y agoI think they mentioned pickling because that's what the multiprocessing queue uses by default. I think Pipe isn't necessarily a drop in replacement depending on the complexity of object you want to share but I have found it significantly faster for simple things.
- while_true_ 4y agoExactly. My assumption is Pipe() is faster for short and simple data because I assume it doesn't do the serialization conversion that Pickle does.
- nijave 4y agoYup, if you need to serialize a complex object manually before piping then you end up paying the price again (like this the multiprocessing Queue) I'm guessing that's part of the reason the article didn't mention it (it looks like they're talking about a Pandas DataFrame which I would say is non-trivial--compared to a primitive type) I'd think Pipe + Parquet should beat filesystem though. That really depends on storage I guess Iirc I played a bit with msgpack and orjson to see if there was anything to gain over Pickle but I don't think it made much difference. You'd probably need to deal with structs Looking at CPython source (3.10), on Windows, you always get a NamedPipe. On other platforms, you get a OS pipe when duplex=False otherwise a socket (socket.socketpair)