3 ms·
One of the lead Arrow developers here (https://github.com/wesm https://github.com/wesm). It's a little bit disappointing for me to see the Arrow project scrutin
by wesm 9y ago
One of the lead Arrow developers here (https://github.com/wesm https://github.com/wesm). It's a little bit disappointing for me to see the Arrow project scrutinized through the one-dimensional lens of columnar storage for database systems -- i.e. considering Arrow to be an alternative technology (i.e. part of the same category of technologies) to Parquet and ORC.
The reality (at least from my perspective, which is more informed by the data science and non-analytic database world) is that Arrow is a new, category defining technology. So you could choose to use it as a main memory storage format as an alternative for Parquet/ORC if you wanted, but that would be one possible use case for the technology, not the primary one.
What's missing from the article is the role and relationship between runtime memory formats and storage formats, and the costs of serialization and data interchange, particularly between processes and analytics runtimes. There are also some small factual errors about Arrow in the article (for example, Arrow record batches are not limited to 64K in length).
I will try to write a lengthier blog post on wesmckinney.com when I can going deeper into these topics to help provide some color for onlookers.
- indogooner 9y agoThanks wesm for clearing this. My understanding (albeit limited) of Arrow was that it would complement Parquet. An example would be speeding up the in-memory representation of Parquet files. I believe mitigating the serialization costs will also help projects like PySpark. Looking forward to your post.
- deleted 9y ago[deleted]
- Thrymr 9y agoIn fact, the conclusion of the article is more positive than the headline makes it out to be: "Therefore, it makes sense to keep Arrow and Parquet/ORC as separate projects, while also continuing to maintain tight integration."
- makmanalp 9y agoI think he buried the lede a bit with the title - his answer being "yes, we should have a separate format for this." The way he phrased his points was a bit odd and seemed inimical at times even though in the end it wasn't - e.g. he mentioned the X100 paper and his C-store compression paper which both talk about lightweight in-memory compression schemes which would suit Arrow's use case well, but then he goes back and says "Arrow probably won't support gzip" (which is much more heavyweight but offers better compression ratio and is more suitable for disk based storage formats) - OK, so that's fine and to be expected then? It turns out, yep, it's expected. His main idea I took away was dispelling the notion that we should have instead been putting effort into slightly repurposing an existing format for our in-memory data layout. It's definitely exciting to see the data science world and the databases world finally interacting a lot more - I think each has a lot to learn from the other. Battle tested techniques on one side, and entirely new use cases to deal with on the other. --- For the curious: Abadi is a well-known name in the databases community, especially with regards to column stores. A few papers that he's co-authored that I like: "Column-Stores vs. Row-Stores: How Different Are They Really?" - speaks about how a different style of query execution is a big part of what drives column store performance, and not just the memory layout itself: http://db.csail.mit.edu/projects/cstore/abadi-sigmod08.pdf http://db.csail.mit.edu/projects/cstore/abadi-sigmod08.pdf And then there's the magnum opus "The Design and Implementation of Modern Column-Oriented Database Systems" which is a huge survey into the subject: https://stratos.seas.harvard.edu/publications/design-and-implementation-modern-column-oriented-database-systems https://stratos.seas.harvard.edu/publications/design-and-imp...
- jacquesnadeau 9y agoYou nailed it. The most exciting part about all of this is being able to move between a "data science" context and a "database" context (and back again) without pain or penalty.
- pg_is_a_butt 9y agoyou suck, and your project sucks. you're all idiots.
- deleted 9y ago[deleted]
- filereaper 9y agoI was under the impression that Apache Arrow was supposed to be the unified storage representation of data that can be used across Apache projects. i.e HBase stores data as Apache Arrow, which can be directly queried by Spark or Hive without the need for serialization. As in eliminating the current overhead of going from HBase HFiles to Spark's internal RDD/DataFrame representation with lots of ser/deser going on on both fronts. I think the comparisons to Parquet and ORC are surface level, I expected Arrow to become the "One storage to rule-them-all". Did I get this wrong? Kindly correct if so.
- jacquesnadeau 9y agoArrow is all about in-memory, not long-term persistence. Systems can write it to disk but it is the one in-memory representation to rule them all, not storage/disk. Disk has its own requirements and challenges outside the scope of Arrow.
- wesm 9y agoPosted here: http://wesmckinney.com/blog/arrow-columnar-abadi/ http://wesmckinney.com/blog/arrow-columnar-abadi/