4 ms·
Note: Vaex has its own memory model. If you input Arrow, it converts to the Vaex data representation. Details here: https://github.com/vaexio/vaex/blob/master/
by wesm 8y ago
Note: Vaex has its own memory model. If you input Arrow, it converts to the Vaex data representation. Details here:
https://github.com/vaexio/vaex/blob/master/packages/vaex-arrow/vaex_arrow/convert.py https://github.com/vaexio/vaex/blob/master/packages/vaex-arr...
One of the primary objectives of Apache Arrow is to have a common data representation for computational systems, and avoid serialization / conversions altogether.
- maartenbreddels 8y agoThat is not correct, I just refer to the buffers/memory, 0 copying going on. Vaex is not really opinionated about the memory model actually. The only exception is the bitmasks that are being copied for now because of an incompatibility with numpy. But if I get a 50GB Arrow dataset, vaex leaves the structure intact. Thanks for your work on Arrow, I hope to support and contribute more to it in the future.
- wesm 8y agoI'm looking at the code I linked, and you are serializing in the general case, it is not zero copy. Unpacking a bitmap is not free.
- maartenbreddels 8y agothe 'convert' name is misleading perhaps, maybe we can agree the proof is in the execution time https://youtu.be/TlTcQJPUL3M?t=478 https://youtu.be/TlTcQJPUL3M?t=478 Anyway, let us celebrate a wider adoption of Arrow! :)