7 ms·
Apache Arrow 4.0
- phissenschaft 5y agoGreat to see Ballista in arrow https://github.com/apache/arrow/pull/9723 https://github.com/apache/arrow/pull/9723
- deleted 5y ago[deleted]
- michael_j_ward 5y agoJust a heads up - ballista has been donated to apache arrow, but they also broke the Rust arrow libraries into separate repos [0][1] in a break from the mono-repo model. Read more at the announcement[3] or the merge request[4]. [0] https://github.com/apache/arrow-datafusion https://github.com/apache/arrow-datafusion [1] https://github.com/apache/arrow-rs https://github.com/apache/arrow-rs [2] https://arrow.apache.org/blog/2021/05/04/rust-dev-workflow/ https://arrow.apache.org/blog/2021/05/04/rust-dev-workflow/ [3] https://github.com/apache/arrow/pull/10096 https://github.com/apache/arrow/pull/10096
- poorman 5y agoLove the progress being made!
- mkoubaa 5y agoReally great to see how fast this thing took flight!
- skrebbel 5y agoI don't understand much about this apache/java data streaming ecosystem (ETL, Kafka, Cassandra, they're all buzzword bingo to me and i don't know what it all means), but maybe someone here can translate this to simpler application programmer terms? I read the overview, and I'm not sure yet, but is this like an in-memory database that runs inside your process? Like, sqlite without disk persistence, or Erlang ETS, but then columnar? I can't completely tell from the overview whether it's about the data format or the querying capability. A columnar ETS alternative would be splendid indeed!
- cookguyruffles 5y agoI'm similarly confused. It seems to be a family of table encodings that sacrifice encoding simplicity for compactness. PyArrow implements parquet files for example, but also I think feather? PR for this project is a mess. Front page should be a bullet point list of deliverables rather than aspirations nobody understands
- vletal 5y agoCome on, it's on the top of the front page https://arrow.apache.org/ https://arrow.apache.org/
- cookguyruffles 5y agoYou don't understand, I'm already a user of PyArrow and it doesn't match that page at all. It handles Parquet and Feather, right? They don't appear on the home page. Clicking the "specifications" link instead starts to talk about Flatbuffers. What's going on?
- papercrane 5y agoArrow is the in-memory format, PyArrow supports loading and saving that data as Parquet and Feather formatted files.
- nerdponx 5y agoParquet and Feather are on-disk file formats. Arrow is an in-memory format. Parquet is not Arrow, but they work well together, in that one can easily be (de)serialized to the other. Feather uses the Arrow IPC format internally.
- Tostino 5y agoTheir FAQ page at least answered some of that for me. Have the same feeling. http://arrow.apache.org/faq http://arrow.apache.org/faq
- vletal 5y ago
- yazaddaruvala 5y agoIs anyone familiar enough to know if Arrow is also targeting usage by libraries like Lucene?
- innagadadavida 5y agoLucene among other things implements an inverted index - basically tracks frequency of a word across different documents. In my opinion, it is already in a columnar like format and highly optimized for the use case and won't see any benefit from changing on disk formats.
- yazaddaruvala 5y ago> it is already in a columnar like format and highly optimized for the use case I mean it has more complex data structures (FST) than just columnar, but yeah for doc values and such I agree, that is exactly why I'm curious if Arrow is targeting that usecase, and will be competitively "highly optimized". > and won't see any benefit from changing on disk formats. I'm less interested in Lucene actually migrating to Arrow (although if it reduces tech-debt they should look into it), I'm most interested in if Arrow will help future Lucene-like libraries get implemented with competitive performance. Also since Arrow version=X is cross-language compatible, it would be amazing to be able to create "Lucene" indexes (or segments) in Java (perhaps for legacy reasons), then use Rust or Go to query the data.
- oregontechninja 5y agoIt's a binary data format, supporting trees, tables, lists, and even blobs. Never used it, I already have sqlite.
- vlmutolo 5y agoTwo major and significant differences from sqlite: No persistent storage. Arrow is meant to be used for in-memory queries. Column-major storage. This enables more efficient data-science-like queries, such as univariate statistics on columns.
- xbar 5y agoHere's a neat story about how an Apple M1 Macbook enjoyed 3x the performance compared to an Apple Intel Macbook using a (hassle to compile) Apache Arrow test instantiation. https://uwekorn.com/2021/01/11/apache-arrow-on-the-apple-m1.html https://uwekorn.com/2021/01/11/apache-arrow-on-the-apple-m1....
- dmitrykoval 5y agoGood progress overall, especially on the Rust side, I wonder if C++ and Rust would at some point follow the same roadmap when it comes higher-level compute features or rather deviate and develop at their own pace. Special kudos to the Rust team for Parquet predicates pushdown feature.
- rubatuga 5y agoFinally has ARM builds for pyarrow!
- liminal 5y agoIs Arrow good for text data or does the columnar format lose its benefits when dealing with lots of arbitrary length strings?